Commit Graph

423 Commits

Author SHA1 Message Date
Nicky Leach cc42a67e7e
fix(adapter-utils): extend the duplex fail-closed run disposition to the CLI lane (#11966)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip runs agents through adapter execution lanes
> - Duplex adapters can lose their control channel before a process
completes
> - The ACP lane already fails closed, but the CLI lane can report false
success
> - This pull request applies the same completion rule to the CLI lane
and shares the loss code
> - The benefit is consistent failure reporting when a duplex channel
closes during a run

## Linked Issues or Issue Description

**What happened?**

A CLI-lane duplex run can lose its control channel before clean process
completion. The run can then report `succeeded` with exit code 0 and no
error code.

**Expected behavior**

The execution target must fail closed when the channel dies before clean
completion. It must return exit code 1, the typed `duplex_channel_lost`
error code, and a short stderr note.

**Steps to reproduce**

1. Start a duplex adapter run through the CLI execution lane.
2. Close the duplex control channel before the process completes
cleanly.
3. Inspect the run result and error code.

**Paperclip version or commit**

Commit `5e01523d4eb6df4a20a0bddd05374c9c42225203`.

**Deployment mode**

Built from source.

**Installation method**

Built from source with pnpm.

**Agent adapter(s) involved**

Claude Code, Codex, Cursor, Gemini, Kimi, OpenCode, and Pi local
adapters.

**Database mode**

Not database-related.

## What Changed

- Add an optional `errorCode` field to `RunProcessResult`.
- Add a one-read completion seam to the execution target process
options.
- Fail closed when a duplex channel dies before clean process
completion.
- Add `settleRunDisposition()` to atomically read and mark orderly
completion.
- Share the typed duplex loss error code across the ACP and CLI lanes.
- Mark non-success terminal results as orderly completion before
teardown.
- Wire the seam through the seven duplex adapters.
- Add regression tests for channel loss, clean completion, and non-clean
terminal results.

## Verification

- `npx vitest run
packages/adapter-utils/src/execution-target-sandbox.test.ts` — 118
passed.
- `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts
-t "sandbox duplex run-disposition seam"` — 4 passed.
- The author confirmed a clean type-check for
`@paperclipai/adapter-utils` and the seven duplex adapter packages.
- Pre-existing environment failures remain outside this change. They
include `EACCES mkdir '/srv/paperclip'` and remote file-size setup
failures.

## Risks

The change alters terminal status for CLI duplex runs that lose control
before clean completion. The typed error code and stderr note keep the
failure visible. The broker marks failed, cancelled, and timed-out
results as orderly completion to prevent false loss events during
teardown.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution, with the standard
GPT-5 context window. The model assisted with the implementation and
test work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-22 11:12:51 -07:00
Nicky Leach 10d2781a29
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox adapters provide controlled execution for untrusted provider
environments.
> - The sandbox channel needs one persistent duplex transport with
strict host control.
> - The transport must remain off unless the instance setting and
provider capability both allow it.
> - The host must detect loss, bound resource use, and expose only safe
telemetry.
> - This pull request adds the broker, gated selection, kill-switch
wiring, fixed observability, and real-process proof.
> - The benefit is safer sandbox execution with bounded failure behavior
and inspectable transport results.

## Linked Issues or Issue Description

No public issue exists for this change. The related pull requests are
#11738 and #11750.

**Problem or motivation**

The sandbox duplex channel needs a host-controlled broker, strict
transport gates, bounded provider input, and safe loss telemetry.
Without these controls, a provider can cause replay, resource growth,
unsafe endpoint selection, or data exposure through telemetry.

**Proposed solution**

Add a host broker with nested time limits, request limits, one-shot
loss, and per-id deduplication. Select duplex transport only when the
instance setting and provider capability both equal true. Assign the
endpoint and nonce on the host. Reject invalid readiness data and use
the file bridge on failure. Add fixed redacted telemetry and a
real-process end-to-end test harness.

**Alternatives considered**

Keep the file bridge as the only transport. This avoids new channel
behavior but does not provide persistent duplex operation for supported
sandbox providers.

**Roadmap alignment**

This change supports the Cloud / Sandbox agents section in ROADMAP.md.

## What Changed

- Add the duplex bridge broker with bounded forward, response, and
gateway wait budgets.
- Bound concurrent requests, lifetime requests, and request-id bytes
before retention or forwarding.
- Select duplex transport only when both required gates are true.
- Assign the loopback port and nonce on the host and enforce a
liveness-only READY frame.
- Fall back to the file bridge after invalid readiness, contamination,
bind failure, or timeout.
- Carry the kill switch through the server, acpx engine, and six local
adapters.
- Add fixed, redacted duplex telemetry with a provider allowlist.
- Add a real-process end-to-end harness for readiness, round trips,
loss, and teardown.
- Add regression coverage for limits, loss, UTF-8 splits, concurrency,
and telemetry dimensions.

## Verification

- Adapter-utils, server, and Daytona typechecks pass locally.
- Adapter-utils tests pass, including the codec, broker,
execution-target sandbox, and real-process harness.
- Server kill-switch tests pass.
- Live Daytona tests pass with the required provider key and skip
without that key.
- The root pnpm-lock.yaml file has no diff.
- The branch contains ten commits after origin/master.

## Risks

- Duplex transport remains disabled unless both gates equal true.
- A provider remains an untrusted boundary and needs least-privilege
credentials and quotas.
- The server telemetry recorder stays deferred; the default recorder
does nothing.
- A provider that pre-binds the host port causes a fail-closed fallback
to the file bridge.
- The change adds no database migration and changes no root lockfile.

## Model Used

OpenAI GPT-5, exact model family GPT-5, large context window, reasoning,
and tool use. The model assisted with Git handoff validation and PR
preparation. The implementation commits came from the engineering
worktree.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: # / Closes #
/ Refs # OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs)
- [x] My branch name describes the change (e.g. docs/... or fix/...) and
contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
- [x] I searched the GitHub PR list for similar PRs and confirmed this
is not a duplicate

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-22 09:01:31 -07:00
Devin Foley fbd20b28d3
fix(grok-local): stop defaulting --permission-mode to dontAsk (#11898)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The `grok_local` adapter runs the native Grok Build CLI in headless
mode for unattended agent heartbeats
> - Grok CLI 1.0 started to enforce the `dontAsk` permission mode as
deny-by-default, and it takes precedence over `--always-approve`
> - The adapter passes both flags on every run, so each run dies on its
first tool call and is still recorded as a success
> - This pull request removes the `dontAsk` default so unattended runs
rely on `--always-approve` alone
> - The benefit is that `grok_local` agents can execute tools again on
current Grok CLI releases

## Linked Issues or Issue Description

No public issue exists. Description per the bug template:

**What happened?**

Every `grok_local` run on Grok CLI 1.0.x stops on its first tool call.
The stream shows the tool call move from `pending` to `failed` with
"User cancelled the execution for tool `run_terminal_command`", and the
session ends with `stopReason: "cancelled"` after one turn. The CLI
exits 0, so Paperclip records the run as succeeded with no work done,
and the issue lands in missing-disposition recovery.

**Expected behavior**

Unattended runs must auto-approve tool executions. The adapter already
passes `--always-approve` for this.

**Steps to reproduce**

In a clean Linux environment with Grok CLI 1.0.3 and `XAI_API_KEY` set,
run the adapter's exact invocation shape:

`grok --output-format streaming-json --permission-mode dontAsk
--always-approve --disable-web-search --single "Run the shell command:
echo ok"`

The tool call is denied. Drop `--permission-mode dontAsk` (or use
`--permission-mode bypassPermissions`) and the same command executes the
tool. On Grok 0.2.x the original combination worked because the CLI
accepted `dontAsk` without enforcing it; the 0.2.39 embedded docs state
the flag takes effect only for `bypassPermissions` / always-approve.

**Paperclip version or commit**

master (917d2350f)

## What Changed

- `packages/adapters/grok-local/src/server/execute.ts`: `permissionMode`
no longer defaults to `dontAsk`. The adapter passes no
`--permission-mode` flag unless one is explicitly configured.
`--always-approve` (default on) remains the unattended policy.
- `packages/adapters/grok-local/src/index.ts`: config doc updated to
explain the new default and the Grok 1.0 semantics.
- `packages/adapters/grok-local/src/server/execute.test.ts`:
default-args assertion now requires the absence of `--permission-mode`;
new test covers explicit `permissionMode` pass-through.

## Verification

- `npx vitest run packages/adapters/grok-local` — 7 files, 29 tests, all
pass.
- `pnpm --filter @paperclipai/adapter-grok-local typecheck` — clean.
- Live matrix against Grok CLI 1.0.3 in a clean sandbox: `dontAsk
--always-approve` denies the first tool call; `--always-approve` alone
executes it; `bypassPermissions --always-approve` executes it; `dontAsk`
alone denies it.

## Risks

- Low risk. Operators who explicitly set `permissionMode` keep their
value verbatim. Only the implicit default changes, and the old default
is what breaks every run on current Grok CLI releases.
- On Grok 0.2.x the flag was unenforced, so omitting it does not change
behavior there.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking, tool use, via
Claude Code CLI.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-21 14:50:43 -07:00
Devin Foley adfbe2d4b9
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Display of the platform-managed default environment across the UI.

**Subsystem affected**

UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.

**Current behavior**

The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."

**Proposed behavior**

Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."

**Reason and benefit**

The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.

**Breaking changes**

None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.

## What Changed

- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.

## Verification

- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).

## Risks

Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.

## Model Used

- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-21 12:51:22 -07:00
Nicky Leach 5bc6031f79
fix(server,ui,claude-local): verify auth on the adapter Test lane and enforce managed-sandbox tenant binding (#11810)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Adapter Test checks whether an agent adapter can run with its
configured environment, and every local-driver adapter (Claude, Codex,
Gemini, OpenCode, Pi, Cursor, etc.) shares this Test route and its UI
resolution logic
> - The Claude ACP Test lane could report pass without checking local or
remote authentication, and the shared Test route and UI had gaps in
environment binding, probe safety, and managed-sandbox resolution that
affect every adapter that uses the Test button, not only Claude
> - This pull request verifies authentication on every Claude ACP
target, and closes the shared Test-route/UI gaps: tenant-binding on the
route, a managed-sandbox-only redirect that matches the real run path,
and a three-tier environment resolution in the UI
> - The benefit is a truthful Test result with safer probe execution and
tenant isolation, for Claude specifically and for every other local
adapter that shares this Test surface

## Linked Issues or Issue Description

**What happened?**

The Claude ACP Test lane returned `status: "pass"` without checking
authentication for some local and non-sandbox targets. Separately, the
shared `/companies/:companyId/adapters/:type/test-environment` route —
used by every local-driver adapter, not only Claude — accepted a foreign
environment id, and its UI resolution did not mirror the server's
managed-sandbox-only redirect.

**Expected behavior**

The Test lane checks the resolved credential and hello probe for every
Claude ACP target. The shared adapter Test route rejects a foreign
environment before it reveals environment details or starts a lease, for
any adapter type. The Test's environment resolution (UI and server)
matches the real run's three-tier resolution, including the
managed-sandbox-only redirect.

**Steps to reproduce**

1. Run the Claude ACP Test lane against a local target without a valid
credential.
2. Run the adapter Test route with an environment id from another
company (any adapter type).
3. Observe the pass result on step 1, or the missing tenant-binding
rejection on step 2.

**Paperclip version or commit**

`933749e01f74e82ce5d315c071be534d04e01158`

**Deployment mode**

Local dev (`pnpm dev`) and server route tests.

**Agent adapter(s) involved**

Claude Code directly (the ACP auth-verification work). The
tenant-binding guard, managed-sandbox-only redirect, and UI three-tier
resolution apply to the shared adapter Test route and affect every
local-driver adapter (Codex, Gemini, OpenCode, Pi, Cursor, etc.), not
only Claude — see "What Changed" below for the split between Claude-only
and shared changes.

**Database mode**

Not database-related.

**Access context**

Both board and agent paths use the affected Test surface, for every
local-driver adapter.

**Additional context**

Two commits that were previously bundled into this PR — a
`plugin-worker-manager` duplex-channel frame-bound fix and a
`workspace-runtime` exit-persist crash fix — are unrelated to the
adapter Test lane and have been split out into their own PRs: #11860 and
#11861.

## What Changed

Claude-only (`packages/adapters/claude-local`):

- Verify `CLAUDE_CODE_OAUTH_TOKEN` and run the hello probe for every
Claude ACP target.
- Keep `adapter_auth_missing` sandbox-only and report missing
non-sandbox credentials as a warning.
- Add a deny-by-default probe environment builder for the ACP and CLI
local probes.
- Log only fixed probe context and allowlisted classifications.
- Seed the host OAuth token into the hello probe environment.

Shared, cross-adapter (`server/src/routes/agents.ts`,
`ui/src/lib/adapter-test-environment.ts`,
`ui/src/components/AgentConfigForm.tsx`,
`ui/src/components/OnboardingWizard.tsx`):

- Add a company-binding guard and a binding assertion for the generic
`/companies/:companyId/adapters/:type/test-environment` route, so a
foreign-company environment id is rejected before any secret resolution
or sandbox lease, for every adapter type.
- Resolve all three server environment tiers (agent default, instance
default, local default) in the UI, and add the managed-sandbox-only
redirect so the Test probes the same target a real run would use.
- Enforce onboarding Test results: block hire on a failed environment
test.

- Add regression tests for authentication, tenant binding, probe safety,
diagnostics, and UI resolution.

## Verification

- Adapter suites pass for the Claude local server probe, remote, ACP,
auth, probe environment, and config paths.
- Server route tests pass, including the five tenant-binding cases.
- UI adapter Test environment resolver tests pass for all three
resolution tiers.
- Adapter package `tsc --noEmit` exits 0.
- Full CI must pass on this pull request.

## Risks

The probe environment now denies caller variables by default. A required
variable that is not on the allowlist could stop a probe from starting.
The route now rejects foreign environment ids with a fixed 403 response.
The managed-sandbox-only redirect changes where the Test (and the login
affordance) probes for every local-driver adapter under that policy, not
only Claude — operators running other local adapters under
managed-sandbox-only will see their Test target move from local to the
managed sandbox, matching what real runs already do. The change limits
secret and diagnostic exposure.

## Model Used

Original implementation: OpenAI Codex, GPT-5; exact context window not
exposed in that run; tool use and code execution.

This revision (commit split and title/description correction): Claude,
Sonnet 5 (claude-sonnet-5). The original title and description described
this PR as Claude-only; review found it also changes the shared adapter
Test route and UI resolution used by every local-driver adapter, and
carried two unrelated server fixes. Claude split those two commits into
#11860 and #11861 via `git rebase --onto` (verified byte-identical to
the original tree minus those commits) and rewrote this description to
reflect the actual scope. No functional code in this PR was authored by
Claude.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-21 11:30:49 -07:00
Nicky Leach 1746783d40
fix(kimi): align package with Node 24 policy (#11890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip now requires Node.js 24.11.0 or later
> - Each workspace package must publish the same Node.js engine
requirement
> - The Kimi adapter entered `master` after the Node.js upgrade branch
started
> - Its package still used Node.js 22 types and had no engine
requirement
> - This pull request aligns the Kimi adapter with the repository
Node.js policy
> - The benefit is that the Node.js policy check passes again on
`master`

## Linked Issues or Issue Description

**What happened?**

The `pnpm check:node-version` command fails on `master`. The Kimi
adapter uses `@types/node` 22 and has no `engines.node` value.

**Expected behavior**

All workspace packages must use Node.js 24 types and declare Node.js
24.11.0 as the minimum version.

**Steps to reproduce**

1. Check out commit `a7e689b3c`.
2. Use Node.js 24.11.0.
3. Run `pnpm check:node-version`.

**Paperclip version or commit**

`a7e689b3c`

**Deployment mode**

Local dev (`pnpm dev`).

## What Changed

- Update the Kimi adapter to use `@types/node` 24.
- Add the repository minimum Node.js engine requirement to the Kimi
adapter package.
- Keep `pnpm-lock.yaml` out of this pull request.

## Verification

- `npx -y -p node@24.11.0 -c 'node --version && pnpm check:node-version
&& pnpm --filter @paperclipai/adapter-kimi-local typecheck'`
- The command reports Node.js `v24.11.0`.
- The Node.js policy check passes.
- The Kimi adapter typecheck passes.
- A broader local suite was started and stopped at the maintainer's
request after the focused checks passed.

## Risks

- Low risk. This change updates package metadata and development types
only.
- The lockfile refresh runs in separate repository automation.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-5, with repository inspection, shell tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-21 10:55:58 -07:00
Nicky Leach 38d8f37172
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip runs across the CLI, server, adapters, plugins, CI, and
container images.
> - These surfaces declared different Node.js versions from 20 through
24.
> - A newer `@types/node` major can expose APIs that the supported
runtime does not provide.
> - Node.js 20 is no longer a suitable project baseline, and Node.js 24
is the current LTS line.
> - This pull request sets Node.js 24.11.0 as one repository-wide
baseline, adds a drift check, and gives users actionable startup
guidance when their runtime is too old.
> - The benefit is one clear runtime contract for development, release,
installation, and published packages.

## Linked Issues or Issue Description

Refs #2734

Refs #11727

Refs #739

## What Changed

- Require Node.js 24.11.0 or newer in all 42 package manifests and
runtime checks.
- Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox
setup, portable installs, and esbuild targets.
- Align every direct `@types/node` declaration on `^24.0.0`.
- Prevent Dependabot from opening major `@types/node` upgrades without a
matching runtime decision.
- Add `.nvmrc` and a CI policy check for Node version drift.
- Update ACP version gates, tests, and user documentation for the new
minimum.
- Print a non-blocking warning on CLI and server startup when Node is
unsupported, with remediation through a version manager or the
documented downloaded `install.sh` workflow.
- Deduplicate that warning when `paperclipai run` boots the CLI and
server in the same process.

## Verification

- `node scripts/check-node-version-policy.mjs`
- `node --check scripts/check-node-version-policy.mjs`
- `node --check cli/esbuild.config.mjs`
- `node --check scripts/generate-npm-package-json.mjs`
- `bash -n scripts/install.sh scripts/test-install-sh-docker.sh
scripts/e2e-install-lifecycle.sh`
- Parsed all 42 package manifests and confirmed `engines.node` is
`>=24.11.0`.
- `git diff --check`
- `vitest run
packages/adapter-utils/src/sandbox-install-command.test.ts` passed with
3 tests.
- `vitest run cli/src/node-version.test.ts` passed with 4 tests.
- Directly exercised the shared warning helper for unsupported-version
messaging and same-process deduplication.
- The focused exe.dev suite could not resolve the locally unbuilt plugin
SDK from this isolated worktree. A full offline workspace install was
also blocked because the package-manager signature verifier requires
registry access. The full suite was not run locally; draft CI performs a
clean install and evaluates the wider impact.

## Risks

- This is a breaking runtime change for users, plugins, and deployments
that still use Node.js 20 or 22.
- Published workspace packages will now produce an engine warning or
failure in strict package managers on older Node.js releases.
- Node.js 24 can reveal dependency, native module, Playwright, or agent
CLI compatibility issues in CI.
- The bootstrap installer now installs Node.js 24 when the current
runtime is older than 24.11.0.
- The portable sandbox fallback is pinned to Node.js 24.11.0 and depends
on that upstream tarball remaining available.
- Unsupported runtimes continue booting after a warning, so a later
incompatibility can still fail at its point of use.
- The CLI and server share the warning policy through the published
`@paperclipai/shared` package; packaging checks must keep that subpath
export available.
- This PR does not commit `pnpm-lock.yaml` because repository policy
assigns lockfile generation to CI.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex based on GPT-5. The exact deployment ID and context
window are not exposed in this session. Reasoning, repository tools,
shell execution, and GitHub tools were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-21 10:17:52 -07:00
David V 233c12f029
feat: add kimi-local adapter for Kimi Code CLI (CLI + ACP engines) (#9967)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Local agent adapters (`claude_local`, `gemini_local`, `grok_local`,
…) are the integration surface that lets Paperclip run coding CLIs on
the host machine
> - The Kimi Code CLI (`kimi`, Moonshot AI) has a documented
non-interactive mode, `kimi -p --output-format stream-json` with session
resume via `kimi -r`, but Paperclip has no built-in adapter for it
> - So Kimi users (especially Kimi membership / OAuth subscribers)
cannot onboard their CLI to Paperclip agent teams
> - This pull request adds a complete built-in `kimi_local` adapter
(both execution engines, session management, instructions + skills
delivery, thinking-effort control, environment test, UI and CLI modules,
docs) following the established `gemini_local`/`grok_local` package
pattern
> - Kimi Code ships an ACP server (`kimi acp`), so the adapter runs on
Paperclip's shared acpx engine by default (streaming transcript with
live tool status, like `claude_local`/`gemini_local`) and falls back to
a headless CLI lane (`kimi -p --output-format stream-json`) when ACP
prerequisites are unavailable
> - The benefit is that Kimi Code becomes a first-class Paperclip agent
lane: selectable in the UI, resumable across heartbeats, with the same
operating context (instruction bundle, skills, effort) and streaming
transcript the other local adapters get

## Linked Issues or Issue Description

- Supersedes #9880 (same branch; expanded from the CLI-only lane into a
complete adapter with the default ACP engine lane, control-plane skill
install, and live transcript wiring)
- Refs #9879 (adapter request for Kimi Code CLI, filed with this PR)
- Refs #163 (original Kimi support request)

Duplicate/related prior PRs, per the dedup search (both appear stale: no
updates or maintainer review since May 2026, and both target an older
Kimi CLI interface; calling them out for reviewer context per
CONTRIBUTING.md):

- Refs #6276 (`feat: add kimi-local adapter`): targets an older
array-based content format (`{type: think}`/`{type: text}` blocks), not
the current documented stream-json schema
- Refs #5202 (`feat(adapter): add Kimi CLI local adapter with Wire
protocol support`): builds on a `--wire` JSON-RPC interface that current
Kimi Code CLI (0.27.0) no longer documents; the current documented
headless interface is `-p --output-format stream-json`

This PR is a fresh implementation against current master and the
currently documented/verified Kimi CLI behavior (see Verification).
Happy to fold in anything useful from the earlier attempts if a reviewer
prefers.

## What Changed

- **New adapter package** `packages/adapters/kimi-local`
(`@paperclipai/adapter-kimi-local`), modeled on
`gemini-local`/`grok-local`:
- `src/server/execute.ts`: spawns `kimi -p <prompt> --output-format
stream-json` (argv array, no shell), `-m <model>` only when configured,
`-r <sessionId>` when the stored session cwd matches the run cwd,
automatic fresh-session retry on unrecoverable-session errors,
headless-safe env (`CI=1`, `NO_COLOR=1`, `KIMI_CODE_NO_AUTO_UPDATE=1`,
`TERM=dumb`; user-configured values win), full remote (ssh/sandbox)
execution lane with runtime install via `@moonshot-ai/kimi-code`
- **Instruction bundle delivery**: the prompt path directive now names
the sibling instruction files (`./HEARTBEAT.md`, `./SOUL.md`,
`./TOOLS.md`) alongside the prepended entry file, and local runs pass
`--add-dir <instructions-dir>` so Kimi can actually open them (matching
`claude_local`). Without this, only the entry file reached Kimi and
agents improvised the operating workflow that `HEARTBEAT.md` documents
- **Thinking effort**: a configured `effort` is forwarded as the
`KIMI_MODEL_THINKING_EFFORT` operational override (Kimi has no
per-invocation effort flag). It is only sent for models that advertise
`support_efforts` (currently `kimi-code/k3`) to avoid provider
rejections, and `medium` maps to `high` since Kimi has no medium tier
(`low`/`high`/`max` pass through)
- **Skills delivery**: desired Paperclip skills are delivered via Kimi's
`--skills-dir` flag from a dedicated per-run directory (a local
snapshot, or the synced snapshot on remote targets), so skills load
reliably and in isolation. Paperclip never overwrites the shared
`$KIMI_CODE_HOME/skills` home, so skills installed by the operator or
other agents are left intact. `--skills-dir` is only passed when at
least one skill is desired, so unconfigured agents keep Kimi's default
skill discovery
- **Live run status**: the adapter now forwards each streamed
stream-json line to `onEvent` (assistant `content` as an assistant
snippet, `tool_calls` as tool-name events), which drives the
issue-thread activity indicator (`currentToolName` /
`lastAssistantSnippet` / `lastEventAt`). Previously the adapter only
wrote the raw run log, so the issue thread showed a stale "no output for
N s" line with no tool or reasoning context while Kimi worked. Tool
results are omitted so the last meaningful "Using X" / snippet is not
overwritten by a generic label
- `src/server/parse.ts`: parses the verified Kimi stream-json event
shapes (`assistant` text, `assistant.tool_calls` with JSON-string
arguments, `tool` results, trailing `meta.session.resume_hint` for
session-id capture) plus failure classifiers (`kimi_auth_required`,
transient network, unrecoverable session). A signaled exit (null exit
code, not a timeout) is now reported as a failure rather than coalesced
to success, and the error message names the terminating signal
- `src/server/skills.ts`: lists/syncs Paperclip skills for the adapter's
skill-management surface
- `src/server/test.ts`: environment test covering CLI resolution + `kimi
--version`, cwd check, auth detection (OAuth credential dirs, keyed
`[providers.*]` in config.toml, or the `KIMI_MODEL_NAME` +
`KIMI_MODEL_API_KEY` env pair), and a live hello probe
- `src/ui/` (stdout-line parser for transcripts, config builder) and
`src/cli/` (stream event formatter) modules
- Root metadata: three managed model aliases
(`kimi-code/kimi-for-coding`, `kimi-code/kimi-for-coding-highspeed`,
`kimi-code/k3`), effort-capable-model metadata (`EFFORT_CAPABLE_MODELS`,
effort mapping helpers), `agentConfigurationDoc`
- Tests: 101 tests across parse, execute (args building, resume gating,
retry, auth error code, timeout, signaled-exit failure, effort
forwarding/gating/mapping, `--add-dir` instructions directive,
`--skills-dir` gating, `onEvent` runtime-event forwarding), ACP engine
(engine resolution, acpx config build, node-version gate), ACP
transcript delegation, environment test, UI parse/build-config
- **ACP engine lane (default)** (`src/server/acp.ts` + shared
`adapter-utils/acpx-engine`): Kimi Code ships an ACP server (`kimi
acp`), so `kimi_local` now runs on Paperclip's shared acpx engine by
default, matching `claude_local`/`codex_local`/`gemini_local`. The
issue-thread transcript streams live (assistant text deltas, tool calls
with a `pending`->`completed` status lifecycle) instead of the CLI
lane's bursty complete-message output. Registered `kimi_local -> "kimi"`
in `ACPX_ADAPTER_AGENT_IDS` and resolved the built-in agent command to
`kimi acp`; `execute.ts` dispatches to the ACP executor first with an
automatic CLI fallback when ACP prerequisites fail (`engine=acp`
requires ACP, `engine=cli` pins the headless lane); `index.ts` falls
back to the shared acpx session codec; the UI/CLI delegate `acpx.*`
events to the shared acpx transcript parser and event formatter. The
headless CLI lane (above) remains as the fallback
- **Registration** (one entry each, mirroring existing adapters): server
adapter registry + `BUILTIN_ADAPTER_TYPES`, `AGENT_ADAPTER_TYPES`
(shared), UI adapter registry + display registry (`Kimi Code`, Moon
icon) + capabilities defaults, CLI adapter registry, `Dockerfile`
(package copy + `npm install --global @moonshot-ai/kimi-code@latest`),
`vitest.config.ts` workspace, `scripts/release-package-manifest.json`
- **Behavioral sets** mirroring `gemini_local` (Kimi resumes sessions
the same way): `GIT_SENSITIVE_LOCAL_ADAPTER_TYPES`,
`SESSIONED_LOCAL_ADAPTERS` (heartbeat + recovery),
`REMOTE_MANAGED_ADAPTERS`, ssh/sandbox execution-target allow-lists,
`ADAPTER_DEFAULT_RULES_BY_TYPE` (`timeoutSec: 0`, `graceSec: 15`), and
`LEGACY_SESSIONED_ADAPTER_TYPES` + `ADAPTER_SESSION_MANAGEMENT` in
adapter-utils
- **UI touch-points**: New Agent default-model branch, AgentConfigForm
command map (`kimi_local: "kimi"`) + model defaults + a Kimi-specific
thinking-effort option list (`Low`/`High`/`Max`, reflecting Kimi's tiers
rather than borrowing Claude's), OnboardingWizard (command map, model
default, `kimi login` / `KIMI_MODEL_NAME + KIMI_MODEL_API_KEY` auth
hints, manual-debug command line), InviteLanding enabled adapters
- **Control-plane skill install** (`cli/src/commands/client/agent.ts`):
`paperclipai agent local-cli` seeded the Paperclip control-plane skills
into `~/.codex/skills` and `~/.claude/skills` so Codex/Claude agents
auto-discover the API reference every run. Kimi had no equivalent
target, so `kimi_local` agents began each session without the
control-plane skill and rediscovered routes (e.g. the company-scoped
`POST /api/companies/{companyId}/issues`) by trial and error. Added
`~/.kimi-code/skills` (honoring `KIMI_CODE_HOME`) as a third install
target for parity. Independent of the per-run `--skills-dir` delivery,
which only applies to explicitly configured skills.
- **Docs**: `docs/adapters/kimi-local.md` (prerequisites, auth options,
config fields including `effort`, session resume, instruction bundle,
skills delivery, control-plane skill install) + a row in
`docs/adapters/overview.md`

Out of scope (deliberately): model profiles, built-in agent
`allowedAdapterTypes` additions.

## Verification\n\nCurrent-master rebase verification (OpenAI Codex,
2026-08-03): 13 focused files / 231 tests pass; adapter-utils, server,
UI, CLI, and Kimi adapter typechecks pass; full repository build and UI
token gates pass. The branch is conflict-free against master at head
`1249df117c5e12e5771b9a570a6340866450619e`.\n\nAutomated (all from repo
root, pnpm 9.15.4, Node 22):

- `vitest run packages/adapters/kimi-local`: 89/89 pass (includes
coverage for the instruction `--add-dir` directive, effort
forwarding/gating/mapping, `--skills-dir` gating, the signaled-exit
failure path, and `onEvent` runtime-event forwarding with cross-chunk
line buffering)
- `vitest run server/src/__tests__/adapter-registry.test.ts
server/src/__tests__/adapter-routes.test.ts
server/src/services/heartbeat-stop-metadata.test.ts
ui/src/adapters/adapter-display-registry.test.ts`: 37/37 pass
- `vitest run cli/src/__tests__/skills.test.ts`: 13/13 pass (the
control-plane skill install target follows the existing Codex/Claude
install path, whose symlink logic is unchanged)
- `vitest run packages/shared`: 307/307 pass; `vitest run
packages/adapter-utils`: pass except one pre-existing, unrelated failure
(`mcp-isolation.integration.test.ts` requires Claude CLI ≥ 2.1.207; host
has 2.1.185, fails identically on unmodified master)
- `pnpm --filter @paperclipai/adapter-kimi-local typecheck|build`, plus
typecheck of `server`, `ui`, `cli`, `adapter-utils`: all clean
- `pnpm install --frozen-lockfile`: passes (the PR diff itself contains
no lockfile changes, per repo policy; verified against a locally
regenerated lockfile)
- `node scripts/check-no-git-push.mjs` and `node
scripts/check-forbidden-tokens.mjs`: pass
- CI note: the `policy` job's release-bootstrap step is expected to stay
red until a maintainer bootstraps the first npm publish of
`@paperclipai/adapter-kimi-local`; see the CI Note for Maintainers
comment. All other contributor-actionable checks are green.

Manual end-to-end (real Kimi CLI 0.27.0, OAuth login, dev server on an
isolated instance):

1. Server `GET /api/adapters` lists `kimi_local` as builtin with correct
capability flags; models endpoint returns the three Kimi models
2. `POST .../adapters/kimi_local/test-environment`: all checks pass,
including a live `kimi -p` hello probe
3. Created a `kimi_local` agent and invoked two heartbeats: run 1
spawned `kimi -p ... --output-format stream-json`, Kimi used its `Read`
tool, produced the expected answer, and the session id was captured from
the `session.resume_hint` meta event; run 2 resumed the **same** Kimi
session (`sessionIdBefore == sessionIdAfter`) via `-r`
4. UI: adapter appears in the New Agent dropdown; selecting it shows the
Kimi command placeholder, the three models, and the Kimi config fields;
the run transcript renders Kimi tool calls via the adapter's stdout
parser

The instruction-bundle, thinking-effort, and `--skills-dir` changes
landed after the manual run above. They are covered by the unit tests
listed under Automated, and the Kimi CLI flags they rely on
(`--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were
confirmed against the installed Kimi Code CLI 0.27.0 (`kimi --help`,
config-file thinking-effort docs).

Screenshots (assets branch on the fork, not part of the diff):

![Kimi Code in the Add a new agent runtime
picker](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/00-kimicode.png)

![Adapter dropdown with Kimi
Code](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/01-adapter-dropdown-kimi.png)

![Kimi adapter selected: command, model, config
fields](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/02-kimi-adapter-selected.png)

![Kimi models in the model
dropdown](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/03-kimi-model-dropdown.png)

![Successful resumed heartbeat run (kimi_local invocation + parsed
transcript)](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/04-successful-resumed-run.png)

![Agents list showing the Kimi Code
label](https://raw.githubusercontent.com/hawikk/paperclip/assets/kimi-local-pr-screenshots/shots/05-agents-list.png)

## Risks

- Low risk to existing behavior: the change is additive, one new
workspace package plus single-entry registrations alongside existing
adapters; no existing adapter code paths are modified.
- The adapter invokes the locally installed `kimi` CLI; like other local
adapters, run behavior depends on the host's Kimi version. The parser is
written against the documented/verified 0.27.0 stream-json schema and
degrades gracefully (malformed lines are skipped, failures surface as
run errors).
- `--skills-dir` overrides Kimi's auto-discovery of user and project
skills for the run. This is intentional (paperclip-managed agents get a
reproducible, isolated skill set), and it is only passed when at least
one Paperclip skill is desired, so unconfigured agents keep default
discovery.
- Thinking effort is only forwarded to models that advertise
`support_efforts` (currently `kimi-code/k3`); `EFFORT_CAPABLE_MODELS`
must be extended when more Kimi models gain support, otherwise a
configured effort is silently ignored for them.
- `Dockerfile` now installs `@moonshot-ai/kimi-code@latest` globally
alongside the other agent CLIs, so image size increases slightly.
- Maintainer action needed for the npm bootstrap gate: the `policy`
job's release-bootstrap step fails until the first npm publish of
`@paperclipai/adapter-kimi-local` (the gate from #5146 that every new
adapter package has passed through). Enrollment with `publishFromCi:
true` is required by the manifest validator (dropping the entry,
`false`, or `private` are all rejected), so this is intentionally left
to a maintainer. Remaining CI lanes are expected to run once it is done.

## Model Used\n\n- **Current-master rebase, conflict adaptation, and
registry-parity coverage:** OpenAI, **GPT-5 Codex** (Codex agent; exact
serving model ID and context-window size were not exposed to the
runtime), with repository, shell, Git, and GitHub tooling. It preserved
Hawik’s commit authorship, reconciled ACPX and environment-capability
changes, added current registry tests, and ran the verification
above.\n- **Adapter implementation and initial review:** Moonshot AI,
**Kimi K3 Coding** (latest), via **Kimi Code CLI v0.27.0**
(`kimi-code/k3` alias, 1M-token context window, thinking mode, agentic
tool use). The CLI agent explored the repo, wrote the adapter
implementation (delegated to a coder sub-agent of the same model), ran
tests, and drafted the first version of this PR body. A second
model-driven review pass (read-only, same model) audited the diff for
security/correctness before submission; its findings (shell-quoting
hardening, auth-detection false positive, session-compaction
registration, test gaps) were fixed and are included.
- **Harness-context fixes and review responses:** Anthropic, **Claude
Opus 4.8** (`claude-opus-4-8`) via Claude Code. Diagnosed from run logs
that Kimi received only the entry instructions file (not the
`HEARTBEAT.md`/`SOUL.md`/`TOOLS.md` bundle) and that `effort` was never
wired, then implemented the instruction `--add-dir` delivery,
`KIMI_MODEL_THINKING_EFFORT` forwarding, and `--skills-dir` skill
delivery, added the accompanying tests and docs, and addressed the
automated review comments (preserving external skills on remote sync,
treating a signaled exit as a failure). Also extended the `paperclipai
agent local-cli` installer to seed the control-plane skills into
`~/.kimi-code/skills` for Codex/Claude parity, wired `onEvent` runtime
events so the issue-thread activity indicator reflects Kimi's tool and
reasoning output live, and built the ACP engine lane (`kimi acp` via the
shared acpx engine, default) so the transcript streams with live tool
status like the other ACP adapters. The Kimi CLI flags, subcommand, and
env var relied on here were verified against the installed Kimi Code CLI
0.27.0.
- All CLI behaviors claimed here (`-p`, `--output-format stream-json`,
`-r` resume, event shapes, `--add-dir`, `--skills-dir`,
`KIMI_MODEL_THINKING_EFFORT`) were verified empirically against the
installed Kimi CLI, not assumed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green *(only the release-bootstrap step
remains red, pending the maintainer npm publish described in Risks)*
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
*(will address all Greptile comments as they arrive)*
- [x] I will address all Greptile and reviewer comments before
requesting merge





---

## Maintainer Addendum (2026-08-20)

The shared acpx-engine and issue-chat changes (run-summary segmentation,
placeholder tool-event coalescing,
`ISSUE_CHAT_TRANSCRIPT_MAX_VISIBLE_ENTRIES` 30 → 400, live-reasoning UI)
have been **extracted to #11761** so the cross-adapter behavior changes
review and revert independently — both commits there preserve @hawikk's
authorship. This PR is now the kimi-specific adapter only (60 files,
+3,793/−8, essentially pure addition); the only shared-engine touch left
is the `kimi acp` command resolution. `publishFromCi` is `true` — the
package name is bootstrapped on npm.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Devin Foley <devin@paperclip.ing>
2026-08-20 12:06:33 -07:00
dependabot[bot] ed1db310a7
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.66.0 to 0.69.0 (#11728)
Bumps
[@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp)
from 0.66.0 to 0.69.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@​agentclientprotocol/claude-agent-acp's
releases</a>.</em></p>
<blockquote>
<h2>v0.69.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.68.0...v0.69.0">0.69.0</a>
(2026-08-16)</h2>
<h3>Features</h3>
<ul>
<li>report changed files to AIR (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>)
(<a
href="450d6b19dc">450d6b1</a>)</li>
</ul>
<h2>v0.68.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.67.0...v0.68.0">0.68.0</a>
(2026-08-14)</h2>
<h3>Features</h3>
<ul>
<li>align typed session failures with AIR protocol (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>)
(<a
href="0581b9cf39">0581b9c</a>)</li>
</ul>
<h2>v0.67.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.67.0">0.67.0</a>
(2026-08-14)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Update to
<code>@​anthropic-ai/claude-agent-sdk</code> v0.3.232 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>)
(<a
href="de0d0e2b7d">de0d0e2</a>)</li>
<li>expose typed session failures for AIR (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/979">#979</a>)
(<a
href="8157ee113e">8157ee1</a>)</li>
<li>publish the model fallback as a warning advisory (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>)
(<a
href="35aaddb3d5">35aaddb</a>)</li>
<li>surface resolved model name in default model option description (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>)
(<a
href="ec73cd8560">ec73cd8</a>)</li>
<li>surface Skill tool calls with name and kind in _meta (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>)
(<a
href="1f09e9a3ca">1f09e9a</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>preserve task plans across prompts (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/974">#974</a>)
(<a
href="1afa940a2c">1afa940</a>)</li>
<li>show a pending title while Claude prepares a file (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/978">#978</a>)
(<a
href="3df1ede89f">3df1ede</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@​agentclientprotocol/claude-agent-acp's
changelog</a>.</em></p>
<blockquote>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.68.0...v0.69.0">0.69.0</a>
(2026-08-16)</h2>
<h3>Features</h3>
<ul>
<li>report changed files to AIR (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>)
(<a
href="450d6b19dc">450d6b1</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.67.0...v0.68.0">0.68.0</a>
(2026-08-14)</h2>
<h3>Features</h3>
<ul>
<li>align typed session failures with AIR protocol (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>)
(<a
href="0581b9cf39">0581b9c</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.67.0">0.67.0</a>
(2026-08-14)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Update to
<code>@​anthropic-ai/claude-agent-sdk</code> v0.3.232 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>)
(<a
href="de0d0e2b7d">de0d0e2</a>)</li>
<li>expose typed session failures for AIR (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/979">#979</a>)
(<a
href="8157ee113e">8157ee1</a>)</li>
<li>publish the model fallback as a warning advisory (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>)
(<a
href="35aaddb3d5">35aaddb</a>)</li>
<li>surface resolved model name in default model option description (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>)
(<a
href="ec73cd8560">ec73cd8</a>)</li>
<li>surface Skill tool calls with name and kind in _meta (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>)
(<a
href="1f09e9a3ca">1f09e9a</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>preserve task plans across prompts (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/974">#974</a>)
(<a
href="1afa940a2c">1afa940</a>)</li>
<li>show a pending title while Claude prepares a file (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/978">#978</a>)
(<a
href="3df1ede89f">3df1ede</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="59a7e9367b"><code>59a7e93</code></a>
chore(main): release 0.69.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1007">#1007</a>)</li>
<li><a
href="450d6b19dc"><code>450d6b1</code></a>
feat: report changed files to AIR (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>)</li>
<li><a
href="5de5d4a2c8"><code>5de5d4a</code></a>
chore(main): release 0.68.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1000">#1000</a>)</li>
<li><a
href="0581b9cf39"><code>0581b9c</code></a>
feat: align typed session failures with AIR protocol (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>)</li>
<li><a
href="11dd73b871"><code>11dd73b</code></a>
Restore native provider state after overrides (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/996">#996</a>)</li>
<li><a
href="e4dba808ea"><code>e4dba80</code></a>
chore(main): release 0.67.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/975">#975</a>)</li>
<li><a
href="de0d0e2b7d"><code>de0d0e2</code></a>
feat(deps): Update to <code>@​anthropic-ai/claude-agent-sdk</code>
v0.3.232 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>)</li>
<li><a
href="35aaddb3d5"><code>35aaddb</code></a>
feat: publish the model fallback as a warning advisory (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>)</li>
<li><a
href="1f09e9a3ca"><code>1f09e9a</code></a>
feat: surface Skill tool calls with name and kind in _meta (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>)</li>
<li><a
href="ec73cd8560"><code>ec73cd8</code></a>
feat: surface resolved model name in default model option description
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.69.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-20 08:40:49 -07:00
Nicky Leach e0b64529b3
feat(auth): normalize agent login in the sandbox onto one session table and a capability contract (#11730)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Sandbox agents need a safe login path for each supported adapter
> - Codex device login and Claude setup-token login used separate
session stores and route logic
> - Separate stores made session lookup, expiry, and login capability
checks harder to keep consistent
> - This pull request unifies both flows on one session table and one
capability contract
> - The benefit is one company-scoped login model with public session
identifiers and shared lifecycle rules

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting (multiple of the above)

**Problem or motivation**

Codex and Claude sandbox login used separate session stores and
different route paths. This split increased the risk of inconsistent
company scoping, session lookup, and cleanup.

**Proposed solution**

Use `adapter_auth_sessions` for both login flows. Use public session
identifiers for API access. Select login behavior from projected adapter
capability data. Share the route spine, lease arguments, runner
lifecycle, and reaper rules.

**Alternatives considered**

Keep two session tables and add matching fixes to both routes. This
keeps duplicate logic and does not provide one capability contract, so
this pull request uses shared infrastructure.

**Roadmap alignment**

This change supports the shipped Cloud / Sandbox agents milestone in
`ROADMAP.md`.

## What Changed

- Unify Codex device login and Claude setup-token login on
`adapter_auth_sessions`.
- Return and look up sessions with company-scoped public session
identifiers.
- Enforce one active session for each company, owner, and adapter.
- Share the login route spine, sandbox lease arguments, runner
lifecycle, and missing-auth check.
- Add a standalone setup-token reaper with adapter-specific row
selection.
- Add optional login capability projection for adapters and drive route
and UI selection from that data.
- Rename the provider flag to `supportsLoginPty` and validate its
deprecated alias.
- Remove the old Claude setup-token session table and add the required
migrations.

## Verification

- Server typecheck passed with `tsc`.
- Database typecheck passed.
- UI typecheck passed with `tsc -b`.
- Codex login service and route suites passed.
- Setup-token session, route, and reaper suites passed.
- Adapter session schema, plugin validator, capability projection, UI
render, and Daytona suites passed.
- GitHub Actions must confirm the complete CI gate after pull request
creation.

## Risks

- The migrations remove short-lived in-flight login rows during
deployment. A login that spans the migration can continue until its
provider lease expires.
- The Codex credential store remains company-scoped. A cross-owner
credential race remains a documented, board-accepted risk.
- API clients that use internal session row identifiers no longer work.
The API accepts only public session identifiers.

## Model Used

Codex, GPT-5, exact runtime model ID not exposed in this handoff, large
context window, reasoning, and repository tool use. The implementing
engineer produced the code with AI assistance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-19 11:51:31 -07:00
dependabot[bot] ec8d7dd21f
build(deps): bump ws from 8.21.1 to 8.21.3 (#11729)
Bumps [ws](https://github.com/websockets/ws) from 8.21.1 to 8.21.3.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/websockets/ws/releases">ws's
releases</a>.</em></p>
<blockquote>
<h2>8.21.3</h2>
<h1>Bug fixes</h1>
<ul>
<li>The server now correctly rejects permessage-deflate offers if the
incoming
<code>client_max_window_bits</code> parameter value is smaller than its
configured
<code>clientMaxWindowBits</code> (e97a20ea).</li>
</ul>
<h2>8.21.2</h2>
<h1>Bug fixes</h1>
<ul>
<li>Fixed a test for <a href="https://github.com/nodejs/citgm">CITGM</a>
(2eb3be0b).</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="c791e707ea"><code>c791e70</code></a>
[dist] 8.21.3</li>
<li><a
href="e97a20eaa6"><code>e97a20e</code></a>
[fix] Reject offers with <code>client_max_window_bits</code> below
config</li>
<li><a
href="787ebf22ce"><code>787ebf2</code></a>
[dist] 8.21.2</li>
<li><a
href="b4d62ebad4"><code>b4d62eb</code></a>
Revert &quot;[ci] Trust Coveralls Homebrew tap&quot;</li>
<li><a
href="e4bb883723"><code>e4bb883</code></a>
[security] Use GitHub PVR as main reporting channel</li>
<li><a
href="2eb3be0bff"><code>2eb3be0</code></a>
[test] Skip test on Node.js versions where it does not apply</li>
<li>See full diff in <a
href="https://github.com/websockets/ws/compare/8.21.1...8.21.3">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=ws&package-manager=npm_and_yarn&previous-version=8.21.1&new-version=8.21.3)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-19 10:30:13 -07:00
dependabot[bot] 536d5880c5
build(deps): bump @cursor/sdk from 1.0.24 to 1.0.28 (#11520)
Bumps [@cursor/sdk](https://github.com/cursor/cursor) from 1.0.24 to
1.0.28.
<details>
<summary>Commits</summary>
<ul>
<li>See full diff in <a
href="https://github.com/cursor/cursor/commits">compare view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=@cursor/sdk&package-manager=npm_and_yarn&previous-version=1.0.24&new-version=1.0.28)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-19 00:48:17 -07:00
Dotta a2bf936f9a
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed worktree services run isolated Paperclip instances with
cloned databases.
> - A reachable service was reported as ready even when its database,
runtime identity, or login path was not usable.
> - The first candidate added verified database seeding and managed
repair in #11665.
> - This pull request consolidates that candidate with signed login
handoff and a complete readiness contract.
> - Post-QA fixes close five defects in repair identity, repair
responses, UI retry, seed journal handling, and seed-source trust.
> - The benefit is a workspace that either opens safely or reports one
accurate recovery action.

## Linked Issues or Issue Description

No public GitHub issue exists for this work, so the problem is described
here.

**What happened**

Managed workspace URLs could return HTTP 200 and report ready while
login failed. QA also found cases where repair used the wrong instance
identity, returned a generic error, left the UI stuck, rejected a safe
journal lag, or trusted a mutable workspace manifest.

**Expected behavior**

Opening a ready workspace signs the board user in to the correct
isolated instance. Provisioning and repair use a registered source and
report a structured recovery state.

**Actual behavior**

Entry depended on a password copied into the clone. Several failure
paths could publish stale readiness, hide the repair precondition, or
trust state that the workspace could modify.

**Additional context**

This pull request includes the commits first published in #11665. That
pull request keeps the original base head for review history. This
consolidated pull request is the merge candidate. Related open readiness
work includes #11575 and #11621.

## What Changed

- Adds a short-lived, signed, single-use login ticket. It binds the
user, workspace, instance, and runtime origin.
- Exchanges the ticket through Better Auth. It creates the session and
cookie through the supported adapter path.
- Adds protected workspace readiness fields for the database, clone
data, login handoff, seed phase, and runtime identity.
- Fails readiness closed when the guest has no company or
execution-workspace binding.
- Binds ticket issuance to the exact cloned user and active company
membership selected for the handoff.
- Verifies every current active board identity through the exact-user
handoff before publication or reuse.
- Gates managed runtime publication on the readiness contract and the
recorded worktree instance identity.
- Refreshes runtime work products from the live runtime row after a port
change.
- Adds one workspace access card with ready, degraded, repairing, and
failed states.
- Uses the runtime response identity for repair. It returns structured
repair precondition errors.
- Lets a valid source journal lag converge during provisioning.
- Binds seed and repair manifests to a source registered outside the
agent-writable worktree.
- Clears recovered UI errors so a successful retry can open the
workspace.
- Makes runtime tests register canonical sources and avoid ports owned
by live host listeners.
- Keeps Vitest on source suites when compiled `dist` trees exist.
- Isolates CLI and adapter tests from ambient AWS and runtime API
environment variables.
- Preserves a 404 response for cross-company workspace ID lookups before
runtime authorization.
- Makes concurrent single-flight coverage independent of
path-canonicalization scheduling order.

## Verification

The following checks passed on the integrated head:

```sh
pnpm -r typecheck
pnpm build
pnpm check:token-gates
pnpm --filter @paperclipai/db check:migrations
```

- The server source lane passed 420 files and 4,953 tests. Five tests
were skipped.
- The CLI lane passed 57 files and 385 tests.
- The database lane passed 26 files and 97 tests.
- The shared package passed 58 files and 506 tests.
- The adapter utility lane passed 640 tests. Four tests were skipped.
- The Claude adapter passed 220 tests. One test was skipped.
- The Codex adapter passed 323 tests.
- The OpenClaw adapter passed 13 tests.
- The OpenCode adapter passed 42 tests.
- The plugin SDK passed 45 tests.
- The workspace runtime suite passed 124 tests.
- The caller-scoped readiness and handoff suite passed 52 tests.
- The workspace provisioning shell suite passed 7 tests.
- The runtime exposure suite passed 17 tests while live host mappings
occupied fixed test ports.
- `git diff --check` passed and the worktree is clean.

The serialized route lane will run in GitHub CI with its normal shards.
No deployment or active-workspace migration was performed.

## Risks

- This is a medium-risk authentication and runtime-readiness change.
- The login ticket uses exact origin, workspace, instance, and user
binding. It has a short expiry and a one-time nonce.
- Runtime publication is stricter. A real readiness, identity, per-user
handoff, or control-plane database disagreement now blocks publication.
- This pull request supersedes #11665 as the merge candidate. Close
#11665 after this pull request merges.
- No new database migration is included. The lockfile and workflow files
are unchanged.
- Deployment and active-workspace migration are intentionally outside
this pull request.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool
use, and code execution produced the main candidate. OpenAI GPT-5
(`gpt-5`) through Codex, with agentic reasoning, tool use, and code
execution, integrated the post-QA fixes and hardened the test gates. The
Codex context-window size was not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-19 02:37:02 -05:00
dependabot[bot] 03ddca4df7
build(deps): bump @agentclientprotocol/codex-acp from 1.1.7 to 1.2.0 (#11515)
Bumps
[@agentclientprotocol/codex-acp](https://github.com/agentclientprotocol/codex-acp)
from 1.1.7 to 1.2.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/codex-acp/releases">@​agentclientprotocol/codex-acp's
releases</a>.</em></p>
<blockquote>
<h2>v1.2.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.14...v1.2.0">1.2.0</a>
(2026-08-11)</h2>
<h3>Features</h3>
<ul>
<li>expose typed session failures for AIR (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>)
(<a
href="54987e1c4a">54987e1</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>normalize cwd filters for Windows sessions (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>)
(<a
href="145ebba5d2">145ebba</a>)</li>
</ul>
<h2>v1.1.14</h2>
<h2>What's Changed</h2>
<ul>
<li>Update codex to 0.147.0 by <a
href="https://github.com/acp-release-bot"><code>@​acp-release-bot</code></a>[bot]
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/375">agentclientprotocol/codex-acp#375</a></li>
<li>feat: support replacing goals through ACP control by <a
href="https://github.com/nikita-ashihmin"><code>@​nikita-ashihmin</code></a>
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/376">agentclientprotocol/codex-acp#376</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.13...v1.1.14">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.13...v1.1.14</a></p>
<h2>v1.1.13</h2>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.12...v1.1.13">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.12...v1.1.13</a></p>
<h2>v1.1.12</h2>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.11...v1.1.12">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.11...v1.1.12</a></p>
<h2>v1.1.11</h2>
<h2>What's Changed</h2>
<ul>
<li>build(deps-dev): bump the npm_and_yarn group across 1 directory with
3 updates by <a
href="https://github.com/dependabot"><code>@​dependabot</code></a>[bot]
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/372">agentclientprotocol/codex-acp#372</a></li>
<li>Fix resuming paused goals by <a
href="https://github.com/nikita-ashihmin"><code>@​nikita-ashihmin</code></a>
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/374">agentclientprotocol/codex-acp#374</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.10...v1.1.11">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.10...v1.1.11</a></p>
<h2>v1.1.10</h2>
<h2>What's Changed</h2>
<ul>
<li>fix: Stop emitting &quot;Conversation interrupted&quot; message by
<a href="https://github.com/Rizzen"><code>@​Rizzen</code></a> in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/358">agentclientprotocol/codex-acp#358</a></li>
<li>Update codex to 0.146.0 by <a
href="https://github.com/acp-release-bot"><code>@​acp-release-bot</code></a>[bot]
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/341">agentclientprotocol/codex-acp#341</a></li>
<li>build(deps): bump the npm_and_yarn group across 1 directory with 2
updates by <a
href="https://github.com/dependabot"><code>@​dependabot</code></a>[bot]
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/362">agentclientprotocol/codex-acp#362</a></li>
<li>feat: support device code authentication via URL elicitation by <a
href="https://github.com/AlexandrSuhinin"><code>@​AlexandrSuhinin</code></a>
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/347">agentclientprotocol/codex-acp#347</a></li>
<li>Update codex to 0.146.1 by <a
href="https://github.com/acp-release-bot"><code>@​acp-release-bot</code></a>[bot]
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/370">agentclientprotocol/codex-acp#370</a></li>
<li>feat: expose provider-neutral ACP goal extension by <a
href="https://github.com/nikita-ashihmin"><code>@​nikita-ashihmin</code></a>
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/371">agentclientprotocol/codex-acp#371</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.9...v1.1.10">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.9...v1.1.10</a></p>
<h2>v1.1.9</h2>
<h2>What's Changed</h2>
<ul>
<li>Throttle ACP plan update snapshots by <a
href="https://github.com/nikita-ashihmin"><code>@​nikita-ashihmin</code></a>
in <a
href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/354">agentclientprotocol/codex-acp#354</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/codex-acp/blob/main/CHANGELOG.md">@​agentclientprotocol/codex-acp's
changelog</a>.</em></p>
<blockquote>
<h2><a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.14...v1.2.0">1.2.0</a>
(2026-08-11)</h2>
<h3>Features</h3>
<ul>
<li>expose typed session failures for AIR (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>)
(<a
href="54987e1c4a">54987e1</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>normalize cwd filters for Windows sessions (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>)
(<a
href="145ebba5d2">145ebba</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="b51bedf600"><code>b51bedf</code></a>
chore(main): release 1.2.0 (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/389">#389</a>)</li>
<li><a
href="2dccf45b53"><code>2dccf45</code></a>
ci: release-please release flow (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/388">#388</a>)</li>
<li><a
href="54987e1c4a"><code>54987e1</code></a>
feat: expose typed session failures for AIR (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>)</li>
<li><a
href="9edc924585"><code>9edc924</code></a>
build(deps-dev): bump hono in the npm_and_yarn group across 1 directory
(<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/380">#380</a>)</li>
<li><a
href="145ebba5d2"><code>145ebba</code></a>
fix: normalize cwd filters for Windows sessions (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>)</li>
<li><a
href="5faefec5d5"><code>5faefec</code></a>
Release v1.1.14</li>
<li><a
href="0d45a13c26"><code>0d45a13</code></a>
feat: support replacing goals through ACP control (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/376">#376</a>)</li>
<li><a
href="a8cedc8d37"><code>a8cedc8</code></a>
Update codex to 0.147.0 (<a
href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/375">#375</a>)</li>
<li><a
href="ea57892f7d"><code>ea57892</code></a>
Release v1.1.13</li>
<li><a
href="91cbfd3046"><code>91cbfd3</code></a>
Release v1.1.12</li>
<li>Additional commits viewable in <a
href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.7...v1.2.0">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=@agentclientprotocol/codex-acp&package-manager=npm_and_yarn&previous-version=1.1.7&new-version=1.2.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-19 00:29:53 -07:00
Harjoth Khara aad97d93fe
fix(hermes): surface real reasoning text from reasoning.available events (#9237)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent runs through the Hermes gateway adapter, its stdout is
parsed line-by-line into transcript entries that the issue chat renders
(the UI fetches the adapter's `./ui-parser` from
`/api/adapters/:type/ui-parser.js` and runs `parseStdoutLine`
client-side)
> - Reasoning-capable models emit a `reasoning.available` gateway event
carrying the model's reasoning text, and the chat renders `thinking`
parts as expandable chain-of-thought
> - The gateway parser mapped `reasoning.available` to a hardcoded
`"Hermes reasoning available"` string and discarded the event payload,
so the "thinking" part had no real content — the indicator looked static
and expanding it revealed nothing (#9209)
> - This pull request extracts the actual reasoning text from the event
payload and uses it as the `thinking` part's text, keeping the old
string only as a fallback for payloads that carry no text
> - The benefit is that the "Hermes reasoning available" indicator now
surfaces the model's real reasoning, which the existing
expandable-thinking UI can display

## Linked Issues or Issue Description

Fixes: #9209

## What Changed

- `packages/adapters/hermes/src/gateway/ui/parse-stdout.ts`: the
`reasoning.available` handler now extracts the reasoning text from the
event `data` via a small helper (`extractReasoningText`), checking the
plausible field names (`reasoning`, `reasoning_text`, `thinking`,
`text`, `summary`, `content`) and recursing one level into nested `data`
/ `payload` records, with ANSI stripped. The prior `"Hermes reasoning
available"` string is kept only as a fallback when no text field is
present.
- `packages/adapters/hermes/gateway-ui-parser.cjs`: applied the
identical logical change to the committed CommonJS mirror (exported as
`./gateway/ui-parser`), keeping the two files in sync.
- `packages/adapters/hermes/src/gateway/ui/parse-stdout.test.ts` (new):
unit tests for the gateway parser (there were none) covering
direct-field, `summary`, nested `data`/`payload` extraction, the no-text
fallback, and regression guards for `message.delta` and plain stdout.

## Verification

Ran from `packages/adapters/hermes`:

- `node_modules/.bin/vitest run src/gateway/ui/parse-stdout.test.ts` →
**8/8 passed**.
- Negative control: stashed the source changes and re-ran the same test
file against the current (pre-patch) parser → **4/8 failed** (exactly
the reasoning-extraction assertions), then restored — confirming the
tests are discriminating, not vacuous.
- `npx tsc --noEmit -p .` → clean.

Real-behavior proof (driving the actual shipped `gateway-ui-parser.cjs`
`parseStdoutLine`) is in the block below.

## Risks

- **Low risk.** Behavior is unchanged for events that carry no
recognizable text field — the `"Hermes reasoning available"` fallback is
preserved (verified). Only the `reasoning.available` branch changed;
`message.delta`, `run.failed`/`run.error`, and the generic/system/stdout
branches are untouched.
- The exact field name in a real `reasoning.available` payload is
defined by the external Hermes gateway and is not present anywhere in
this repo, so the extraction is intentionally defensive across several
plausible field names rather than pinned to one. If the real event nests
the text differently than `data` / `payload`, it will fall back to the
existing placeholder (i.e. no regression vs. today). Happy to tighten
the field list against real gateway traffic if a maintainer can share a
sample.

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`) via Claude Code, with tool use and
local test execution (ran vitest/tsc against the change). Planning, code
review, and the real-behavior proof were done with Claude (Opus 4.8) in
the same session.

## Real behavior proof

**Behavior addressed:** A `reasoning.available` Hermes gateway event now
produces a `thinking` transcript part containing the model's real
reasoning text, instead of a static `"Hermes reasoning available"`
placeholder with no content behind it (#9209).

**Real environment tested:** Drove the actual shipped production
artifact — `packages/adapters/hermes/gateway-ui-parser.cjs`, the exact
module the UI loads via `/api/adapters/hermes-gateway/ui-parser.js` and
runs to parse gateway stdout — on Node v24.16.0, macOS. The input is a
raw stdout line in the exact format emitted by
`packages/adapters/hermes/src/gateway/server/execute.ts`
(`[hermes-gateway:event] run=… event=reasoning.available data=…`). Only
the external gateway boundary (the raw line) is synthesized; the parser
code path is the real one.

**Exact steps or command run after this patch:**
```
# BEFORE = git show HEAD:…/gateway-ui-parser.cjs ; AFTER = patched artifact
node proof.cjs   # requires each parser build and calls parseStdoutLine(line, ts)
# line = [hermes-gateway:event] run=run-abc123 event=reasoning.available \
#        data={"text":"Checking whether the cache key includes the tenant id before I refactor the lookup."}
```

**Evidence after fix:**
```
===== BEFORE (master / old code) =====
[ { "kind": "thinking", "ts": "…", "text": "Hermes reasoning available" } ]
thinking part carries real reasoning text? -> NO (static placeholder, nothing for the UI to expand)

===== AFTER (this patch) =====
[ { "kind": "thinking", "ts": "…",
    "text": "Checking whether the cache key includes the tenant id before I refactor the lookup." } ]
thinking part carries real reasoning text? -> YES
```
Additional cases through the same shipped artifact after the patch:
```
-- nested payload (data.payload.reasoning) --
{"kind":"thinking","ts":"…","text":"Weighing two migration orders."}
-- bare signal, no text field (regression guard) --
{"kind":"thinking","ts":"…","text":"Hermes reasoning available"}     # fallback preserved
-- message.delta still works (regression guard) --
{"kind":"assistant","ts":"…","text":"Hello","delta":true}
```

**Observed result after fix:** The `reasoning.available` event yields a
`thinking` part carrying the model's real reasoning text (top-level or
nested), which the existing expandable-thinking rendering in the chat
can display. Events with no text field still yield the original
placeholder, and unrelated events are unaffected.

**What was not tested:** I did not run against a live Hermes gateway —
Paperclip's Hermes gateway binary and its credentials aren't available
on this machine, and no captured real `reasoning.available` payload
exists in the repo, so the exact wire field name is inferred (hence the
defensive multi-field extraction + safe fallback). I also did not render
the full React chat component in jsdom; the change is confined to the
parser, and the chat's expandable `thinking` rendering already exists
(`ui/src/components/IssueChatThread.tsx`). CI / unit tests here are
supplemental to the runtime proof above.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched `9209 in:body` and keyword variants — none found)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change
(`fix/hermes-reasoning-available-payload`) and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (no
user-facing docs describe this behavior; none needed)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (will confirm once CI runs on the
PR)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(will address on review)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-18 14:19:04 -05:00
Devin Foley 48f4ae16ac
fix(codex-local): keep a promoted device-login credential when re-seeding the managed home (#11578)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The codex_local adapter supports a device login that runs in a
trusted sandbox and promotes the credential into the per-company managed
Codex home
> - The managed-home re-seeding step treats every regular-file
`auth.json` as apikey-mode residue and removes it so the shared-home
symlink can be restored
> - The promotion writes the company credential as a regular file, so
the first environment Test or run after a successful login deletes it
> - On a server with no shared Codex login — any containerized
deployment — nothing replaces the file, and the UI reports that the
sandbox has no ready authentication right after it reported a successful
login
> - This pull request makes the cleanup identity-anchored: a
subscription credential whose identity the shared source does not hold
survives re-seeding
> - The benefit is that a device login stays usable after Test and runs,
on hosts with and without a shared Codex login

## Linked Issues or Issue Description

No public GitHub issue covers this. The problem is described in-PR
following the bug template. Related public PRs:
[#11237](https://github.com/paperclipai/paperclip/pull/11237) added the
sandbox device login and the credential promotion,
[#11097](https://github.com/paperclipai/paperclip/pull/11097) added its
building blocks, and the `ensureSymlink` heal for stale copies came from
the fix for #5028.
[#9621](https://github.com/paperclipai/paperclip/pull/9621) touches the
adjacent sandbox auth sync-back lane but not this defect.

**Subsystem affected**

packages/adapters/codex-local — managed `CODEX_HOME` seeding
(`codex-home.ts`).

**Current behavior**

A successful device login promotes the subscription `auth.json` into the
company Codex home as a regular file, and the UI reports the login as
authenticated. The next `seedManagedCodexHome` call — the environment
Test probe and every execute both run it — removes any regular-file
`auth.json` when no API key is configured, because the cleanup assumes
such a file is apikey-mode residue left by a previous run. It then
symlinks `auth.json` from the shared source home. On a server whose
shared home has no Codex login (a container image, for example), there
is no source to symlink, so the home ends with no credential at all. The
Test probe then reports "The sandbox has no ready authentication for
this adapter" immediately after a successful login, and a fresh login
repeats the same cycle. On a server whose shared home does hold a login,
the symlink silently replaces the promoted account with the host
account.

**Expected behavior**

The credential a device login promoted stays in the company home across
Test probes and runs. The #5028 heal (a stale regular-file copy of the
shared credential becomes a symlink to the live source) and the
apikey-residue cleanup keep working.

**Steps to reproduce**

1. Run the server in an environment whose shared Codex home
(`$CODEX_HOME` or `~/.codex`) has no `auth.json`.
2. Complete a Codex device login for a company; the promotion writes the
company home `auth.json` and the UI reports authenticated.
3. Click Test on a codex_local agent (or start a run). The probe reports
no ready authentication, and the promoted `auth.json` is gone from the
company home.

**Proposed solution**

Make the cleanup identity-anchored, the same rule the promotion and the
cache vend already use. A regular-file `auth.json` survives re-seeding
when it holds a usable subscription identity that the shared source does
not also hold, and the shared symlink does not replace it. A
same-identity regular file is still the #5028 stale copy and is still
healed into the symlink, because the symlink serves the same account
with live, rotating tokens. An apikey-mode or unreadable file is still
removed.

## What Changed

- `seedManagedCodexHome` reads the target `auth.json` before the cleanup
and keeps it when `readSubscriptionAccountId` yields an identity the
shared source `auth.json` does not hold. The kept file is excluded from
the shared symlink pass, and the function logs a fixed line when it
keeps the file.
- The function doc comment states the kept-promoted-credential rule.
- Four new `seedManagedCodexHome` test cases: a promoted credential with
no shared auth, a promoted credential with a different shared identity,
the same-identity #5028 heal, and apikey-mode residue removal.

## Verification

```sh
cd packages/adapters/codex-local
npx tsc --noEmit         # clean
npx vitest run           # 28 files, 321 passed, 1 skipped
```

The four new cases fail on the previous code: the first two observed the
promoted file deleted (and, with a shared login present, replaced by the
shared symlink).

## Risks

Low risk. The change narrows one deletion path. Deployments that never
use the device login see no difference: without a promoted subscription
file, the cleanup and the symlink behave exactly as before, and the
#5028 heal is pinned by an existing test plus a new same-identity test.
The one deliberate behavioral shift: after a device login, the promoted
company credential now stays authoritative over the shared host login
for that company — which is the promotion's documented contract ("the
company credential slot").

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and
code execution — investigation, implementation, and tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-17 20:11:07 -07:00
dmndbrp-oss 7ef75f5636
fix(opencode-local): retry models preflight during transient contention (#9225)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Local CLI adapters are responsible for starting agent runtimes and
validating that their configured models are usable before a run starts.
> - The OpenCode local adapter checks `opencode models` during model
discovery and preflight validation.
> - On hosts with a shared Ollama daemon, that lightweight metadata call
can transiently queue behind an active generation and time out or return
a short failure.
> - Treating that transient contention as a hard adapter failure
prevents otherwise valid local OpenCode runs from starting.
> - This pull request adds a small bounded retry/backoff around OpenCode
model discovery while keeping the existing per-attempt timeout and
surfacing a final failure when retries are exhausted.
> - The benefit is fewer false adapter failures during local Ollama
contention without changing shared Ollama configuration or hiding
genuinely stuck model discovery.

## Linked Issues or Issue Description

No public GitHub issue exists for this adapter reliability bug.

Bug description:
- What happened: `opencode models` can transiently time out or fail
while a shared local Ollama daemon is busy serving another OpenCode
generation, causing the adapter preflight to fail before the actual run
starts.
- Expected behavior: transient model-list contention should be retried
briefly before declaring the adapter unavailable.
- Steps to reproduce: run an OpenCode local adapter using an
Ollama-backed model while another `opencode run` is actively generating
against the same daemon, then trigger model discovery/preflight during
that contention window.
- Paperclip version/commit: observed on the current Paperclip
master-line OpenCode local adapter before this change.
- Deployment mode: local trusted / local CLI adapter execution with a
shared local Ollama daemon.

Related search:
- Searched public GitHub issues for `opencode models preflight retry`;
no matching issue found.
- Searched public GitHub PRs for `opencode models preflight retry`; no
matching PR found. The only search hit was unrelated OpenClaw gateway
authentication work (#6121).

## What Changed

- Added bounded retry/backoff to OpenCode model discovery: three total
attempts with 2s and 4s waits between failures.
- Preserved the existing 20s per-attempt `opencode models` timeout.
- Retry covers timeout and non-zero process exits, while spawn-level
failures still surface immediately.
- Added unit coverage for transient fail -> timeout -> success behavior
and exhausted retry behavior.
- Updated existing OpenCode environment diagnostic tests with explicit
timeouts for the intentional retry/backoff path.

## Verification

- `pnpm --filter @paperclipai/adapter-opencode-local exec vitest run
src/server/models.test.ts src/server/execute.test.ts` -> 2 files passed,
13 tests passed.
- `pnpm --filter @paperclipai/adapter-opencode-local typecheck` ->
passed.
- `pnpm vitest run
server/src/__tests__/opencode-local-adapter-environment.test.ts` -> 1
file passed, 3 tests passed.
- Branch diff against current `upstream/master` is limited to
`packages/adapters/opencode-local/src/server/models.ts`,
`packages/adapters/opencode-local/src/server/models.test.ts`, and
`server/src/__tests__/opencode-local-adapter-environment.test.ts`.

## Risks

Low risk. This only changes OpenCode model discovery behavior and keeps
the preflight bounded. A genuinely unavailable `opencode models` call
still fails after three attempts, and command spawn failures are not
masked.

## Model Used

OpenAI Codex, GPT-5.5 coding agent, tool-enabled repository editing and
shell verification in a local Paperclip workspace.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Test <test@paperclip.ing>
2026-08-17 14:02:30 -07:00
seb-veto 93ff6a8771
fix(cursor-cloud): drop unreachable Paperclip API callback for remote… (#8546)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run via per-adapter execute paths; the `cursor_cloud` adapter
runs the agent in Cursor's cloud (remote), orchestrated server-side via
the Cursor Agent SDK
> - Local adapters receive a run-scoped Paperclip JWT
(`supportsLocalAgentJwt=true`) injected as `PAPERCLIP_API_KEY` so the
agent can call the Paperclip API; `cursor_cloud` is intentionally
`supportsLocalAgentJwt=false` (no JWT minted for a remote worker)
> - But `buildPaperclipEnv` always sets `PAPERCLIP_API_URL` (defaulting
to the local runtime host), so the remote cloud worker is handed a
callback URL it can neither reach nor authenticate against
> - Any agent-initiated Paperclip API call from the cloud worker
therefore fails with a 401 (or is unreachable), producing log noise and
confusing failures
> - This pull request drops the callback wiring when there is no usable
key, so cloud-side Paperclip tools degrade to a clean no-op
> - The benefit is no spurious 401s from remote cloud runs, with run
results unaffected (delivered server-side via the Cursor Agent SDK)

## Linked Issues or Issue Description

No existing public issue — describing the bug inline (per
`.github/ISSUE_TEMPLATE/bug_report.yml`):

**What happened**

`cursor_cloud` runs emit 401s when the remote cloud agent attempts
Paperclip API calls. Root cause: `buildPaperclipEnv`
(`packages/adapter-utils/src/server-utils.ts`) always sets
`PAPERCLIP_API_URL` (local runtime default), while `cursor_cloud` has
`supportsLocalAgentJwt=false`, so no `PAPERCLIP_API_KEY` is minted — URL
present, key absent → 401 / unreachable from `buildWakeEnv` in
`packages/adapters/cursor-cloud/src/server/execute.ts`.

**Expected behavior**

A remote cloud worker that is not issued a run JWT should not attempt
(and fail) Paperclip API callbacks.

**Steps to reproduce**

1. Configure a `cursor_cloud` agent (runs in Cursor's cloud;
`supportsLocalAgentJwt=false`).
2. Trigger a run that causes the cloud agent to make a Paperclip API
call.
3. Observe a 401 (or unreachable) because `PAPERCLIP_API_URL` points at
an unreachable local runtime and no key is present.

**Paperclip version**

Reproduced on current `master` (cutover base `e68188c43`).

**Deployment mode**

Self-hosted control plane, `cursor_cloud` adapter (remote execution in
Cursor's cloud).

**Related PRs (searched; none duplicate this fix):**

- #8197 — `claude_local` opt-out of the sandbox *bridge* for
direct-reachable remote SSH targets. Related family, but the opposite
situation: that path keeps the callback because the remote is reachable
**and** has a run token. `cursor_cloud` has neither, so here the
callback is removed.
- #8130, #4794, #8025 — `PAPERCLIP_API_URL`/loopback injection for
**local** agents (distinct from the remote cloud worker case).
- #401 — alternative agent-auth scheme (run-ID header when no bearer
token); different approach, not overlapping with this targeted fix.

## What Changed

- `packages/adapters/cursor-cloud/src/server/execute.ts`: in
`buildWakeEnv`, when there is no usable `PAPERCLIP_API_KEY`, delete
`PAPERCLIP_API_URL` and `PAPERCLIP_API_BRIDGE_MODE` so the remote worker
performs no Paperclip API callbacks. Informational `PAPERCLIP_*` vars
(run id, agent id, company id, task, wake reason) still flow. When a key
*is* present (operator-provided), the URL is retained.
- `packages/adapters/cursor-cloud/src/server/execute.test.ts`: new test
asserting no callback vars are injected when no run JWT is present;
positive assertion that the URL is retained when a key is present.

## Verification

- `pnpm exec vitest run
packages/adapters/cursor-cloud/src/server/execute.test.ts` → **5/5
pass**.
- `pnpm --filter @paperclipai/adapter-cursor-cloud typecheck` →
**green**.
- Confirmed result delivery does not depend on this callback:
`execute()` reads results server-side via `Agent.getRun()` and
`run.wait()`.

## Risks

- **Low risk.** Only affects the env handed to remote `cursor_cloud`
workers. No schema/migration/behavioral change to result delivery (which
is server-side). When an operator explicitly provides
`PAPERCLIP_API_KEY`, the callback URL is retained, preserving
intentional callback setups.

## Model Used

- **Claude Opus 4.8** (Anthropic), extended/high reasoning mode, via the
Cursor agent with tool use + code execution. Diagnosis grounded in the
adapter/runtime code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change
(`fix/cursor-cloud-skip-unreachable-callback`) and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no documented behavior changes)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending CI run)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending review)
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Sebastian Heyneman <sebastian@joinnova.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-17 14:02:16 -07:00
Nicky Leach c1c46f1e4e
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Claude local adapter supports subscription login through a
sandbox
> - The new-agent page must show login before the user creates an agent
> - Test results must not expose raw sandbox diagnostics or secret
values
> - This pull request adds the login UI to both Test lanes and closes
the diagnostic boundary
> - The branch also adds durable cleanup recovery for failed sandbox
teardown
> - Reusable sandboxes must retain both their recorded teardown
configuration and a valid lifecycle path until destruction succeeds
> - The benefit is a usable login flow with fixed public checks,
redacted server logs, and recoverable sandbox cleanup

## Linked Issues or Issue Description

Related public work:
[#9488](https://github.com/paperclipai/paperclip/pull/9488) adds
first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and
remote runs. Related public issue:
[#2681](https://github.com/paperclipai/paperclip/issues/2681) requests
Claude Code subscription support. This pull request adds the login
transport and new-agent UI flow that those changes do not provide.

**Subsystem affected:** Claude local adapter, server login probes,
sandbox provider setup, cleanup recovery, and the new-agent UI.

**Problem or motivation:** The Test lanes did not show the sandbox login
panel in all supported cases. Test results also exposed raw probe
diagnostics, and JSON escapes could end secret redaction early.

**Proposed solution:** Surface the login capability through the bundled
provider manifest. Prepare the same probe runtime in the ACP lane. Send
diagnostics only to redacted server logs. Keep Test checks on fixed
public messages. Normalize login URL hints to allowlisted HTTPS Claude
and Anthropic hosts. Consume JSON escapes during redaction. Preserve
failed sandbox cleanup state across retries and restarts, and prevent
deletion from severing the lifecycle context of a live reusable sandbox.

**Alternatives considered:** Keep raw diagnostics in Test checks or
trust login URL text from the sandbox. Both choices increase information
exposure. Keep separate probe behavior in the ACP lane. That choice
would leave the two Test lanes inconsistent.

## What Changed

- Surface the sandbox login panel on both Test lanes.
- Reconcile the bundled Daytona plugin manifest so
`supportsSetupTokenLogin` reaches the UI capability gate.
- Prepare the ACP Test lane with the same probe runtime as the CLI Test
lane.
- Add the `claude_acp_login_probe_unavailable` warning when the ACP
probe cannot run.
- Send raw sandbox diagnostics only to redacted server logs.
- Keep Test checks on fixed public messages in the ACP, managed-config,
and CLI paths.
- Normalize login URL hints to allowlisted HTTPS Claude and Anthropic
hosts.
- Redact JSON and escaped-JSON secret values, including escaped quotes
and backslashes.
- Preserve orphan cleanup records across provider failures, restarts,
and unavailable plugins.
- Atomically block environment deletion while a live reusable sandbox
lease still depends on it.
- Verify pending cleanup destroys plugin sandboxes with the provider
configuration recorded on the lease, even after the current environment
configuration changes.

## Verification

- Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`.
- Focused environment route/service/runtime coverage passes: 196 tests
across 3 files.
- `pnpm -r typecheck` passes.
- `pnpm build` passes.
- The full Vitest run completed with 4,754 passing and 28 failing tests.
All 23 source-test failures reproduce unchanged on parent head
`58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate
executions from stale `server/dist` output. The failures are unrelated
macOS path/listener and scheduler-fixture failures, so there is no new
bad commit for bisect to localize.
- All required CI checks pass for the current head, including build,
typecheck/release registry, all server and workspace shards, serialized
server suites, canary, and e2e.
- A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`
reports 5/5, “safe to merge,” with no blocking failure remaining.

## Risks

- A probe or redaction change could hide useful server diagnostics.
- An allowlist change could reject a valid Claude login URL.
- Cleanup recovery changes could affect provider teardown ordering.
- An environment with a live reusable sandbox can no longer be deleted
until the owning issue or execution workspace completes teardown.
- The implementation keeps public Test messages fixed and sends detail
to redacted server logs.

## Model Used

OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code
execution enabled; extended reasoning enabled. The implementation author
used AI-assisted development.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and documented the result
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation or confirmed no separate
documentation change is needed
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-17 13:42:51 -07:00
Nicky Leach e52b8a343f
fix: ACP run lifecycle corrections — failure settlement, workspace sync-back, lease cleanup (#11454)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent adapters run ACP sessions and manage runtime, workspace, and
lease resources.
> - Several failure paths left runtime bridges, staged workspaces, or
environment leases active after an error.
> - These leaks reduce run reliability and can leave later runs without
clean resources.
> - This pull request closes the failure paths, applies one teardown
policy, and adds regression tests.
> - The benefit is consistent failure settlement and safer reuse of
agent workspaces and leases.

## Linked Issues or Issue Description

**What happened?**

ACP runs could leave runtime bridges, staged workspaces, or environment
leases active after failures. Claude and Gemini ACP runs did not restore
the sandbox workspace on teardown. Lease release stopped when one lease
returned an error.

**Expected behavior**

Each ACP failure must return an error result and settle its resources.
Teardown must run each step, release leases independently, and restore
the host workspace when the sandbox ends. Pending cleanup leases must
receive bounded retry attempts.

**Steps to reproduce**

1. Run an ACP session that fails after runtime creation or during turn
preparation.
2. Run an ACP session that fails during a warm hit or staged runtime
handoff.
3. Run lease cleanup with more than one lease when the first release
returns an error.
4. Inspect the result phase, teardown calls, workspace state, and lease
metadata.
5. Run the regression suites listed in the Verification section.

## What Changed

- Settle every ACP failure after runtime creation with an error result
and one sandbox.startup span closure.
- Close the ACP runtime and remove warm entries after every pre-turn
failure.
- Run all teardown steps, record teardown errors, release staging leases
in finally, and prevent duplicate teardown.
- Dispose staged runtimes after seam failures and remove borrowed staged
entries with identity guards.
- Add fail-open workspace sync-back teardown for Claude and Gemini ACP
adapters.
- Isolate lease release errors and add bounded retry sweeps for stranded
pending_cleanup leases.
- Atomically claim pending_cleanup retries and clamp attempt readers to
keep the five-attempt bound.
- Default absent provider reusableLeases values to false and align the
fake provider with its runtime declaration.
- Add regression tests for engine, adapter, server, and shared
environment behavior.

## Verification

- [x] `npx vitest run
packages/adapter-utils/src/acpx-engine/execute.test.ts` — 124 tests
passed.
- [x] `npx vitest run
packages/adapters/codex-local/src/server/acp.test.ts
packages/adapters/claude-local/src/server/acp.test.ts
packages/adapters/gemini-local/src/server/acp.test.ts` — 61 tests
passed.
- [x] `npx vitest run server/src/__tests__/environment-runtime.test.ts
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts
server/src/__tests__/reusable-leases-default.test.ts
server/src/__tests__/environment-routes.test.ts
packages/shared/src/environment-support.test.ts` — passed.
- [x] All listed suites ran from the repository root.
- [x] GitHub CI completed successfully for
`cfc349c9f232711433897915112a1c52c0e462ca`.
- [x] Greptile completed with a 5/5 confidence score and no blocking
finding.

## Risks

The engine changes affect failure settlement and teardown order across
ACP runs. The server changes add retry state to existing lease metadata
without a schema migration. The adapter changes restore workspaces after
sandbox execution. Regression tests cover the changed paths. GitHub CI
and Greptile passed for the current head.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. This change fixes runtime
reliability and does not duplicate a roadmap feature.

## Model Used

OpenAI GPT-5 Codex. The model used tool-based repository inspection,
GitHub operations, and code review support. The runtime does not expose
a context-window value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (for example, `docs/...` or
`fix/...`) and contains no internal Paperclip ticket id or
instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-15 18:52:08 -07:00
LeeJ a53cc8819b
fix(claude-local): pipe print prompt via stdin (#9500)
Fixes #2444.
Refs #4947.

The `claude_local` adapter launched Claude Code as
`claude --print - --output-format stream-json --verbose`. Paperclip writes
the rendered task prompt to Claude's stdin, but current Claude Code releases
can treat the stale `-` positional marker as the prompt itself, so Claude
received the literal string `"-"` instead of the issue body. The customer's
task ran against no content at all.

The fix keeps `--print` mode and stdin delivery, and removes the stale `-`.

Adds regression coverage on both sides of the delivery path: a `claude_local`
assertion that `--print` is present, `"-"` is absent and the prompt still
reaches stdin, and an adapter-utils case proving the sandbox run-log command
wrapper preserves stdin while streaming logs.

Authored by @elJayAdvisor, whose commit is included unchanged with their
authorship. The branch had gone stale and was showing CONFLICTING; the
conflict was in `execution-target-sandbox.test.ts`, where their new test was
added at the same point as master's `creates the process session directories
only in the launch exec` case and git interleaved the two into one hunk.
Resolved by taking master's file and re-inserting their test whole, after
checking every helper it needs still exists there.

Verified: the bug was still live on master at `execute.ts:838`; the
regression test genuinely catches it — restoring the stale `-` fails
`expect(captured.argv).not.toContain("-")`; `@paperclipai/adapter-claude-local`
and `@paperclipai/adapter-utils` typecheck clean; 67 pass across the two test
files. All CI gates green; Greptile 5/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:34:15 -07:00
Devin Foley 9b1fd42ac1
test(grok-local): isolate billing env in usage cost test (#11285)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Local adapters report run output, token use, and cost data.
> - The Grok local adapter now reports real token use and cost data.
> - Its new billing test must prove the no-key path and the API-key
path.
> - The no-key assertion used the caller environment without isolation.
> - This made the test fail when `XAI_API_KEY` was already set.
> - This pull request isolates that environment state in the test.
> - The benefit is stable coverage for the cost gate from #10433.

## Linked Issues or Issue Description

Refs #10433

**What happened?**

The Grok local usage and cost test asserted subscription billing while
it still used the ambient process environment. If `XAI_API_KEY` was set
before the test ran, the adapter selected API billing instead. The
subscription assertion could then fail on a developer machine or a CI
runner with provider credentials.

**Expected behavior**

The test should prove the subscription path with no `XAI_API_KEY`. It
should also prove the API billing path with a test key.

**Steps to reproduce**

1. Start from `master` after #10433.
2. Set `XAI_API_KEY` in the shell environment.
3. Run `vitest` for
`packages/adapters/grok-local/src/server/execute.test.ts`.
4. Observe that the subscription half can take the API billing branch
without test isolation.

**Paperclip version or commit**

`master` after #10433.

**Deployment mode**

Built from source.

## What Changed

- Isolated `XAI_API_KEY` with save, delete, set, and restore logic
around both billing assertions.
- Gave the subscription and API billing checks separate run ids and temp
roots.

## Verification

- `XAI_API_KEY=ambient-test-key corepack pnpm exec vitest run
packages/adapters/grok-local/src/server/execute.test.ts`
- `corepack pnpm --filter @paperclipai/adapter-grok-local typecheck`

## Risks

Low risk. This changes test setup only. It does not change Grok local
adapter runtime behavior.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-5 Codex local coding agent. The agent used shell tools,
GitHub CLI, and local test execution. The context window size was not
exposed in this run.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Claude <noreply@paperclip.ing>
2026-08-13 13:58:05 -07:00
dependabot[bot] 3040db3343
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.63.0 to 0.66.0 (#11314)
Bumps
[@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp)
from 0.63.0 to 0.66.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@​agentclientprotocol/claude-agent-acp's
releases</a>.</em></p>
<blockquote>
<h2>v0.66.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a>
(2026-08-07)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the
minor group (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>)
(<a
href="7f27c47c5c">7f27c47</a>)</li>
<li>expose provider-neutral ACP goal extension (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>)
(<a
href="8b31dea11b">8b31dea</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>publish and replace Claude goals reliably (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>)
(<a
href="f8fd3ab822">f8fd3ab</a>)</li>
</ul>
<h2>v0.65.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a>
(2026-08-05)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>)
(<a
href="b965dd2191">b965dd2</a>)</li>
<li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the
minor group (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>)
(<a
href="15b4eb46f3">15b4eb4</a>)</li>
<li><strong>deps:</strong> Bump <code>@​hono/node-server</code> from
1.19.17 to 2.1.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>)
(<a
href="f9123f3e18">f9123f3</a>)</li>
<li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>)
(<a
href="0988438428">0988438</a>)</li>
<li><strong>steering:</strong> settle a steered turn at idle, not at the
interrupt (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>)
(<a
href="a84b81080a">a84b810</a>)</li>
</ul>
<h2>v0.64.2</h2>
<h2>Bug Fixes</h2>
<ul>
<li>restore the single-tool representation for ExitPlanMode (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>)
(4302a4b)</li>
</ul>
<h2>v0.64.1</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a>
(2026-08-02)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>release 0.65.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>)
(<a
href="0936ec281e">0936ec2</a>)</li>
</ul>
<h2>v0.64.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a>
(2026-07-30)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>)
(<a
href="8e099e8442">8e099e8</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 7 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>)
(<a
href="3f60921959">3f60921</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@​agentclientprotocol/claude-agent-acp's
changelog</a>.</em></p>
<blockquote>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a>
(2026-08-07)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the
minor group (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>)
(<a
href="7f27c47c5c">7f27c47</a>)</li>
<li>expose provider-neutral ACP goal extension (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>)
(<a
href="8b31dea11b">8b31dea</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>publish and replace Claude goals reliably (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>)
(<a
href="f8fd3ab822">f8fd3ab</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a>
(2026-08-05)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>)
(<a
href="b965dd2191">b965dd2</a>)</li>
<li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the
minor group (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>)
(<a
href="15b4eb46f3">15b4eb4</a>)</li>
<li><strong>deps:</strong> Bump <code>@​hono/node-server</code> from
1.19.17 to 2.1.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>)
(<a
href="f9123f3e18">f9123f3</a>)</li>
<li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>)
(<a
href="0988438428">0988438</a>)</li>
<li><strong>steering:</strong> settle a steered turn at idle, not at the
interrupt (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>)
(<a
href="a84b81080a">a84b810</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.1...v0.64.2">0.64.2</a>
(2026-08-02)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>restore the single-tool representation for ExitPlanMode (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>)
(<a
href="4302a4b0b6">4302a4b</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a>
(2026-08-02)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>release 0.65.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>)
(<a
href="0936ec281e">0936ec2</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a>
(2026-07-30)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>)
(<a
href="8e099e8442">8e099e8</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 7 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>)
(<a
href="3f60921959">3f60921</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li><strong>steering:</strong> add opt-in host-owned fallback (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/919">#919</a>)
(<a
href="43af4ec29e">43af4ec</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/903">#903</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="6b405138fc"><code>6b40513</code></a>
chore(main): release 0.66.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/961">#961</a>)</li>
<li><a
href="8aaf608b4e"><code>8aaf608</code></a>
ci: fix release flow (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/971">#971</a>)</li>
<li><a
href="f8fd3ab822"><code>f8fd3ab</code></a>
fix: publish and replace Claude goals reliably (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>)</li>
<li><a
href="133337ffe5"><code>133337f</code></a>
ci: simplify the release flow and make it agent-friendly (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/965">#965</a>)</li>
<li><a
href="8b31dea11b"><code>8b31dea</code></a>
feat: expose provider-neutral ACP goal extension (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>)</li>
<li><a
href="bba912728f"><code>bba9127</code></a>
ci: validate PR titles against release-please conventions (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/962">#962</a>)</li>
<li><a
href="7f27c47c5c"><code>7f27c47</code></a>
feat(deps-dev): Bump globals from 17.8.0 to 17.9.0 in the minor group
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>)</li>
<li><a
href="6d608cb399"><code>6d608cb</code></a>
chore(main): release 0.65.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/957">#957</a>)</li>
<li><a
href="a84b81080a"><code>a84b810</code></a>
feat(steering): settle a steered turn at idle, not at the interrupt (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>)</li>
<li><a
href="b965dd2191"><code>b965dd2</code></a>
feat(deps-dev): Bump nanoid from 3.3.16 to 3.3.17 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.66.0">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=@agentclientprotocol/claude-agent-acp&package-manager=npm_and_yarn&previous-version=0.63.0&new-version=0.66.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-13 10:06:19 -07:00
Philip D'Souza 4660562fde
fix(opencode-local): make the model-availability probe non-fatal (#10294)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents run through adapters; the `opencode-local` adapter shells out
to the OpenCode CLI and, before each run, does a pre-flight `opencode
models` **availability probe** to fail fast on a misconfigured
`provider/model`.
> - That probe was written to **throw on any probe failure** — a
timeout, a non-zero exit, or a transient `Unexpected error` from the CLI
— which aborts the whole heartbeat run.
> - In practice the CLI probe fails transiently (provider hiccup, cold
cache, momentary CLI error). When that happens *after* the agent has
already done its work, the run dies before its terminal disposition is
written, so the platform reopens the issue and re-runs it — a spurious
crash/re-run loop that affects every agent on the OpenCode adapter.
> - This PR makes the probe **non-fatal when it cannot run**: it warns
and proceeds with the configured model, letting the real invocation be
authoritative.
> - It deliberately **keeps** the genuine guard: when the probe
*succeeds* and the configured model is absent from a non-empty list, it
still throws (this is what catches misconfigured slugs).
> - The benefit is that a best-effort pre-flight check can no longer
take down an otherwise-healthy run, while the useful misconfiguration
guard is retained.

## Linked Issues or Issue Description

No public GitHub issue exists; describing inline (bug).

**What happened:** an OpenCode-adapter agent run terminated at the
adapter level with `` `opencode models` failed: Unexpected error ``. The
failure landed after the agent had produced its work, so the
terminal-status update never applied and the run was reopened and
re-executed.

**Expected:** a transient failure of the `opencode models` availability
*probe* should not abort the run — the probe is a best-effort pre-flight
guard, not a gate.

**Actual:** the probe threw on timeout / non-zero exit / empty output,
aborting the run and discarding the completed work + disposition.

**Scope:** both the local (`models.ts`) and remote/SSH (`execute.ts`)
probe paths; affects any agent on the `opencode_local` adapter.

Related PRs (context / prior art):
- Refs #5119 — added the remote execution-target model-probe validation
this PR softens.
- Refs #3291 — closed prior attempt to make the `opencode_local` model
probe non-blocking (at agent-create time; different entry point).
- Refs #8014 — related open work raising the probe timeout (20s → 60s);
complementary, not overlapping.

## What Changed

- `models.ts` (`ensureOpenCodeModelConfiguredAndAvailable`): if
discovery throws (probe can't run) or returns an empty list, **warn and
proceed** with the configured model instead of throwing. The "model
present in a non-empty list" check is unchanged and still throws when
the configured model is genuinely absent.
- `execute.ts` (`ensureRemoteOpenCodeModelConfiguredAndAvailable`):
remote probe **timeout / non-zero exit / empty output** now warn and
return (proceed) instead of throwing. The remote model-absent guard
still throws.
- `models.test.ts`: the local "discovery cannot run" case now asserts
the probe **proceeds** with the configured model (was: asserts it
rejects).
- `execute.test.ts`: added remote regression tests — non-zero exit,
timeout, and empty output all proceed; a successful probe missing the
configured model still rejects.

## Verification

```bash
pnpm --filter @paperclipai/adapter-opencode-local typecheck   # clean
# opencode-local server suite (default 5s per-test timeout is too tight for the
# heavy SSH tests on some machines; use a realistic timeout):
node node_modules/.pnpm/vitest@*/node_modules/vitest/vitest.mjs run \
  packages/adapters/opencode-local/src/server/models.test.ts \
  packages/adapters/opencode-local/src/server/execute.test.ts \
  packages/adapters/opencode-local/src/server/execute.remote.test.ts \
  --testTimeout=45000
```

Result: typecheck clean; all opencode-local server tests pass, including
the new remote fail-open tests and the retained "model unavailable on
the remote target" guard test.

## Risks

- **Fail-open behavior (intentional).** When the probe can't run, a
genuinely misconfigured model is no longer caught at pre-flight — it
surfaces at the real invocation instead. This is the accepted tradeoff:
the probe is best-effort, and the real invocation is authoritative. The
high-value guard (probe succeeds + model absent from a non-empty list)
is retained, so the common misconfiguration — a bad `provider/model`
slug — is still caught.
- No API, schema, or migration changes. Behavior change is confined to
the two probe helpers. Low risk overall.

## Model Used

Anthropic **Claude Opus 4.8** (`claude-opus-4-8`), used via Claude Code
with agentic tool use (repo search, file editing, shell/code execution)
and extended reasoning. Used to diagnose the crash, implement the fix,
and write the tests; the change was reviewed before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change
(`fix/opencode-model-probe-non-fatal`) and contains no internal ticket
id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — N/A
(internal adapter behavior; no user-facing docs affected)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (functional gates:
tests/build/e2e/typecheck/security). Review/Greptile gate re-running
after this update.
- [ ] Greptile is 5/5 with no open P2s — re-triggered after addressing
both P2s (remote test coverage + this template-complete description)
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-12 16:44:03 -07:00
Constantine 2f1c0e011e
fix(hermes): surface silent nonzero exit failures (#10107)
## Thinking Path

- Followed a silent nonzero Hermes exit from child-process result
parsing through heartbeat run, runtime, task-session, and agent
finalization.
- Found two gaps: the adapter could return `errorMessage: null` for a
numeric nonzero exit, and heartbeat later reused the nullable adapter
field instead of its normalized fallback.
- Kept timeout, signal-cancellation, and specific parsed diagnostics
authoritative.

## Linked Issue(s) / Bug Report

Related to #9751 (stderr classification) and #9519 (exit-zero
finalization), but this is a separate failure mode.

Reproduction: run Hermes with a child result equivalent to `exitCode:
1`, `timedOut: false`, and no parsed diagnostic. The heartbeat row
derives `Adapter failed`, while runtime/task-session/agent finalization
can persist null diagnostics.

## What Changed

- Give silent numeric nonzero Hermes exits a stable fallback such as
`Hermes exited with code 1`.
- Preserve specific parsed errors and timeout/signal semantics.
- Reuse the normalized persisted run error for recovered runtime state,
task-session `lastError`, and agent `errorReason`.
- Add adapter-level and embedded-Postgres regressions.

## Verification

- Hermes adapter `execute.onspawn.test.ts` — 7 passed.
- Focused heartbeat normalized-error regression — 1 passed (91 skipped).
- `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` —
passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check origin/master...HEAD` — passed.

Independent review also ran the full recovery file: the changed
regression passed; one unrelated pre-existing timing-sensitive test
timed out.

## Risks / Rollout Notes

Low risk. Fallback text is used only when a numeric nonzero exit has no
better diagnostic. Existing timeout, signal, and parsed-error precedence
remains unchanged.

## Model Used

OpenAI Codex `gpt-5.6-sol` with repository inspection, test execution,
and independent read-only review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (not
applicable: internal diagnostics only)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: cucurigoo <cucurigoo@users.noreply.github.com>
2026-08-12 16:09:35 -07:00
Jannes Stubbemann 1f7959bc69
fix(codex-local): skip benign stderr warnings when deriving the fallback run error (#10003)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents execute through adapters; the codex_local adapter runs the
Codex CLI and reports each run's outcome, including an error message
when the CLI exits nonzero
> - When no error can be parsed from the CLI's JSONL output, `toResult`
in `packages/adapters/codex-local/src/server/execute.ts` falls back to
the first non-empty stderr line as the run error
> - The adapter itself passes the approvals-bypass flag, so the CLI's
first stderr line is always the benign startup warning "YOLO mode is
enabled. All tool calls will be automatically approved."
> - Failed runs therefore record that warning as their error, hiding the
real cause (for example an OpenAI API error further down in stderr) and
making failures hard to diagnose from the run record
> - This pull request derives the fallback error from the first
meaningful stderr line, skipping a conservative set of known benign
lines, and keeps the existing behavior when every line is benign
> - The benefit is that failed Codex runs surface the actual failure
reason instead of a harmless startup warning, without ever producing an
emptier message than before

## Linked Issues or Issue Description

No public issue exists for the codex_local case. The same bug class was
fixed for gemini-local in Refs #5099 and Refs #3476; this PR applies the
equivalent fix to codex_local.

**What happened?**
On a multi-tenant cloud deployment of Paperclip, several codex_local
runs failed and their run records showed `error_code=adapter_failed`
with the error text "YOLO mode is enabled. All tool calls will be
automatically approved." That is a benign Codex CLI startup warning,
printed on every run because the adapter passes the approvals-bypass
flag itself. The real failure (an OpenAI API error printed later in
stderr) was never surfaced.

**Expected behavior**
When the Codex CLI exits nonzero and no error was parsed from its JSONL
output, the run error should be the first stderr line that actually
explains the failure, not a startup warning the adapter itself provoked.

**Steps to reproduce**
1. Configure a codex_local agent and make the underlying Codex CLI
invocation fail after startup (for example, configure a model id the
active credentials cannot use).
2. Run the agent so the CLI exits nonzero with no parsed JSONL error.
3. Inspect the run's error message: it shows the YOLO approvals warning
(the first stderr line) instead of the real error printed further down
in stderr.

## What Changed

- Added `firstMeaningfulStderrLine` next to `firstNonEmptyLine` in
`packages/adapters/codex-local/src/server/execute.ts`, with a
conservative benign-line predicate covering the YOLO approvals warning
and `[paperclip] ...` diagnostic lines the adapter injected (for example
ACP fallback notes).
- Used it only in the `toResult` fallback error derivation. If every
stderr line is benign, the existing chain still applies (first non-empty
line, then `Codex exited with code N`), so the message never gets
emptier than today. Logging is unchanged.
- Added
`packages/adapters/codex-local/src/server/execute.stderr-error.test.ts`:
four end-to-end cases through `execute()` with a mocked CLI process,
plus unit coverage for the new helper. Tests were written first and
confirmed failing before the fix.

## Verification

- `pnpm exec vitest run
packages/adapters/codex-local/src/server/execute.stderr-error.test.ts`
(7 tests pass; 5 failed before the fix as expected)
- `pnpm exec vitest run packages/adapters/codex-local` (21 files, 188
tests pass)
- `pnpm run typecheck` in `packages/adapters/codex-local` (clean)

## Risks

Low risk. Only the derived fallback `errorMessage` changes, and only
when a benign line would otherwise have been picked; parsed JSONL
errors, logging, retry/quota/auth classification inputs, and the
empty-stderr exit-code fallback are untouched. The benign-line list is
deliberately conservative (exact prefixes) so real errors are never
skipped.

## Model Used

Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 16:09:10 -07:00
LeeJ 5521d768b2
fix(claude-local): avoid root-only skip permissions failure (#9463)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Claude local is one of the adapter paths that lets operators run
Claude Code through a local Paperclip runtime.
> - Claude Code rejects `--dangerously-skip-permissions` when the
process is running as root or through sudo.
> - Local/self-hosted Paperclip deployments may run inside root-owned
Docker/runtime processes, so the Claude local adapter can fail before it
reaches the actual runtime/auth condition.
> - Paperclip already uses a curated `--allowedTools` list instead of
`--dangerously-skip-permissions` for remote Claude targets.
> - This pull request applies the same safer permission strategy to
local root processes while preserving existing local non-root and remote
behavior.
> - The benefit is clearer, safer Claude local diagnostics/execution in
containerized setups without widening permissions beyond the existing
explicit tool allowlist.

## Linked Issues or Issue Description

No directly matching public issue or PR found.

Bug description:

- **Problem:** `claude_local` can fail its local probe/execution path
when Paperclip runs from a root-owned local container/runtime because
Claude Code refuses `--dangerously-skip-permissions` under root/sudo.
- **Actual behavior:** The adapter may fail immediately with Claude's
root/sudo guard before validating the real Claude runtime/auth state.
- **Expected behavior:** Local root processes should use the same
explicit allowlist strategy Paperclip already uses for remote targets,
while local non-root behavior remains unchanged.
- **Environment:** Local/self-hosted Docker or container-style Paperclip
runtime where the app process UID is `0`.

Related but different: #4926 covers MCP config propagation for the
Claude local adapter, not the root/sudo permission flag behavior fixed
here.

## What Changed

- Added root-aware permission argument selection for the Claude local
adapter.
- Preserved current local non-root behavior:
`--dangerously-skip-permissions` is still used when allowed.
- Preserved current remote behavior: remote targets continue using
explicit `--allowedTools`.
- Changed local root behavior to use the explicit `--allowedTools` list
instead of `--dangerously-skip-permissions`.
- Threaded process UID awareness through Claude local probe and
execution paths.
- Added unit coverage for skip-disabled, remote, local non-root, local
root, and UID-unavailable behavior.

## Verification

```sh
./node_modules/.bin/vitest run --config g15-vitest-claude-local.config.mjs \
  packages/adapters/claude-local/src/server/permissions.test.ts
```

Result:

```text
1 file passed
8 tests passed
```

```sh
pnpm --filter @paperclipai/adapter-claude-local typecheck
```

Result:

```text
@paperclipai/adapter-claude-local typecheck passed
```

Additional local smoke:

- Ran a disposable root-container Claude adapter diagnostic against this
patch.
- The diagnostic no longer fails with Claude's root/sudo
`--dangerously-skip-permissions` error.
- It proceeds to the actual environment-specific Claude auth/runtime
result.
- No credentials, tokens, hostnames, private paths, or internal
Paperclip issue references are included in this PR.

Public duplicate checks performed:

```sh
gh pr list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools'
gh issue list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools'
```

## Risks

Low-to-medium risk adapter behavior change:

- Local root Claude runs will now use explicit `--allowedTools` rather
than broad skip-permissions behavior.
- That is intentionally safer, but an environment depending on broader
implicit tool access under root may now need the adapter allowlist to
include any required tools.
- Local non-root behavior is unchanged.
- Remote behavior is unchanged.
- No database migrations, API contract changes, or UI changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex `gpt-5.5` via Hermes Agent, with shell/file/tool use for
repository inspection, patching, local verification, and GitHub CLI
operations.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: LeeJ <elJayAdvisor@users.noreply.github.com>
2026-08-12 16:08:14 -07:00
Nicky Leach e31951a17d
feat: Claude agent setup-token login in a sandbox (#11286)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Claude agents that run in a remote sandbox need a safe in-product
login path
> - The existing host login route cannot open a pseudo-terminal inside
that sandbox
> - The login flow must protect the browser code, the login URL, and the
OAuth token at every step
> - This pull request adds the parser, the runner, a Daytona
pseudo-terminal transport, and a guarded, owner-bound session route
behind an injectable transport
> - The route stays inert in the default build and fails closed until a
sandbox provider binds the live transport
> - The benefit is a company-scoped setup-token flow with one-time
secret delivery, redaction, and fail-closed transport checks, ready for
a later staged production rollout

## Linked Issues or Issue Description

**Agent or provider**

Claude Code setup-token login for sandbox agents.

**Why this adapter is useful**

Sandbox agents need a supported way to sign in without host credentials.
An authorized owner completes the browser step and receives the token
one time.

**How the agent is invoked**

When a sandbox provider binds the injectable transport, the server
starts `claude setup-token` through a sandbox pseudo-terminal, sends the
browser code to the matched prompt, and returns the token through the
guarded session route. The default build does not bind the transport. In
that state the start route fails closed with a fixed no-secret `503`. It
does not start a process and it does not hold a sandbox lease.

**Additional context**

The transport is injectable, so each sandbox provider binds its own
pseudo-terminal. This pull request adds the Daytona transport but does
not bind it in the production server. A production wiring needs a lease
manager, a live pseudo-terminal factory, a durable token store, and its
own security review. The route keeps secrets out of logs, activity
details, errors, telemetry, and non-owner responses.

## What Changed

- Add strict parsers for the setup-token URL, the prompt, and the
success token.
- Add a login runner that drives the `claude setup-token` command
through a pseudo-terminal.
- Add the Daytona pseudo-terminal transport and the sandbox plugin
wiring.
- Add a company-scoped, owner-bound login session service with rate
limits, a reaper, cleanup, and one-time token delivery.
- Add the guarded session routes at
`/agents/:id/setup-token-login-sessions/*` behind an injectable
transport. The routes become the live login path only when a provider
binds the transport.
- Keep the start route fail-closed in the default build. It returns a
fixed no-secret `503` and it does not bind `setupTokenLogin`.
- Keep the existing host route `POST /agents/:id/claude-login` in place.
This pull request does not replace it.
- Keep confidential responses behind a fail-closed TLS transport guard
with `Cache-Control: no-store`, and extend redaction for the new fields.
- Export the parser and the runner from the Claude local server entry,
and document the new session routes in the OpenAPI spec.

## Verification

- `pnpm --filter @paperclipai/server exec vitest run setup-token-route
setup-token-session`
- `pnpm --filter @paperclipai/adapter-claude-local exec vitest run`
- `pnpm --filter @paperclipai/server run typecheck`
- Confirm that the pull request checks pass on GitHub.

## Risks

- Low user-facing risk on merge. The default build does not bind the
transport, so the production start route stays fail-closed with a `503`.
The merge does not change the production login behavior.
- When a provider later binds the transport, the flow starts a live
sandbox process and holds a short-lived in-memory secret. Cleanup must
stop the child before it releases the sandbox lease.
- The transport guard fails closed when the deployment does not provide
a trusted TLS path. A wrong proxy allowlist can block a valid request.
- The production wiring is out of scope. It needs a lease manager, a
live pseudo-terminal factory, a durable token store, and its own
security review before the server binds `setupTokenLogin`.

## Model Used

Anthropic Claude Opus 4.8 assisted the implementation. It used extended
reasoning, code execution, repository tool use, and a 200,000-token
context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (the
OpenAPI spec covers the new session routes; no user-facing documentation
needs changes)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 13:02:49 -07:00
Nicky Leach e5a7fd7038
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Agent adapters connect Paperclip to tools such as the Codex command
line tool.
> - A sandboxed Codex agent may start without a credential.
> - The operator needs a safe sign-in flow that does not expose
credentials to the shared package or the sandbox.
> - This pull request adds a company-scoped device-login flow with a
temporary Daytona sandbox.
> - The flow promotes the credential only after readiness checks pass
and removes the temporary sandbox after use.
> - The result lets an operator sign in to a sandboxed Codex agent from
the agent form.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting (multiple of the above)

**Problem or motivation**

A Codex adapter that runs in a sandbox cannot authenticate when the
company has no pre-provisioned Codex credential.

**Proposed solution**

Add a company-scoped device-login session. Start a temporary sandbox,
run `codex login --device-auth`, stream the code and URL, verify
readiness, promote the credential, and delete the sandbox.

**Alternatives considered**

Pre-provisioning a credential does not support first-time sandbox login.
Keeping the credential in the login sandbox does not provide a durable
company credential.

**Roadmap alignment**

This supports the roadmap item for cloud and sandbox agents.

**Additional context**

The flow uses a five-minute cleanup reaper, compare-and-set status
changes, and a PostgreSQL advisory lock to protect promotion and
cleanup.

## What Changed

- Add the adapter login-session contract, database table, and migration.
- Add company-scoped server routes and a service for sandbox device
login.
- Add credential promotion, readiness checks, and cleanup after login.
- Add restart-safe cleanup for abandoned login sandboxes.
- Add sandbox login controls to the agent creation and edit forms.
- Keep device-login and vendor identifiers out of public shared and
adapter UI symbols.

## Verification

- `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`
passed with 310 tests at the submitted commit.
- The server login route, service, and reaper tests passed with 45 tests
at the submitted commit.
- The agent form render tests passed with 26 tests at the submitted
commit.
- The public-symbol leak check passed at the submitted commit.
- A live Daytona sign-in flow still requires confirmation by a user with
a live sandbox.

## Risks

The migration adds a new company-scoped table. A promotion or cleanup
race could remove a credential or leave a sandbox active, so the service
uses claims, compare-and-set transitions, and an advisory lock. The live
Daytona flow needs operator confirmation because local tests do not
provide a real browser sign-in.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution, extended reasoning.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-12 08:58:25 -07:00
Valentin Marchaud b5ebda1dca
fix(grok-local): report real token usage and cost instead of hardcoded zeros (#10433)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Cost/usage tracking is core to that: the dashboard shows per-agent
spend so a team can see what their AI workforce is costing them
> - The `grok_local` adapter (xAI's Grok Build CLI) is a newer adapter
than `claude_local`/`codex_local`, and its usage/cost wiring was left
incomplete
> - Every `grok_local` run persists
`usage.inputTokens/outputTokens/cachedInputTokens = 0` and `costUsd =
null` in `heartbeat_runs`, unconditionally, even though the underlying
`grok` CLI reports real, non-zero token counts and cost per turn in its
own JSON stream
> - This pull request wires the parser to actually read
`usage`/`total_cost_usd` from the CLI's terminal `end` event, threads
those values into the adapter's execution result, and marks them
`usageBasis: "per_run"` so the heartbeat service doesn't incorrectly
delta them against a prior run on a resumed session (matching how
`claude_local`/`codex_local` already do this)
> - The benefit is accurate cost/usage visibility for any self-hosted
Paperclip instance running Grok Build agents, instead of a dashboard
that always reads zero

## Linked Issues or Issue Description

Fixes: #10432

## What Changed

- `packages/adapters/grok-local/src/server/parse.ts`: `parseGrokJsonl()`
now reads `usage.input_tokens` / `usage.output_tokens` /
`usage.cache_read_input_tokens` / `total_cost_usd` from the terminal
`end` event and returns them on `ParsedGrokJsonl` (previously discarded
entirely).
- `packages/adapters/grok-local/src/server/execute.ts`: `toResult()` now
populates `usage.inputTokens/outputTokens/cachedInputTokens` from the
parsed values instead of hardcoded `0`, sets `usageBasis: "per_run"`
(each `--single` invocation reports usage for just that process, not a
running session total), and surfaces `costUsd` only when `billingType
=== "api"` (metered) — subscription/OAuth billing has no marginal dollar
cost, so it stays `null` there, but token counts are populated for both
billing types since usage visibility is useful regardless of billing
model.
- `packages/adapters/grok-local/src/server/parse.test.ts`: added a test
asserting usage/cost extraction from a representative `end` event
payload, and updated the existing exact-equality test for the new
fields.
- `packages/adapters/grok-local/src/server/execute.test.ts`: added a
test covering both subscription billing (tokens populated, `costUsd:
null`) and API-key billing (tokens populated, real `costUsd`), and
asserting `usageBasis: "per_run"` in both cases.

## Verification

- `pnpm vitest run packages/adapters/grok-local/src/server/parse.test.ts
packages/adapters/grok-local/src/server/execute.test.ts` — 9/9 passed
- `tsc --noEmit` on the `grok-local` package — clean
- Verified against a real self-hosted Paperclip instance running `grok`
CLI `0.2.112` with SuperGrok subscription (OAuth) auth: confirmed the
raw CLI stream reports real `usage`/`total_cost_usd` (e.g.
`"usage":{"input_tokens":21560,...},"total_cost_usd":0.0564448`) that
was previously discarded before ever reaching
`heartbeat_runs.usage_json`, which always showed all-zero tokens
regardless of real usage.

## Risks

- Low risk, additive change scoped entirely to the `grok_local`
adapter's usage/cost reporting path — no change to control flow, session
handling, or process execution.
- `usageBasis: "per_run"` mirrors the existing, already-tested pattern
in `claude_local`/`codex_local` execute paths, so the heartbeat
service's per-run vs. session-cumulative delta logic is exercised the
same way.
- `costUsd` is intentionally left `null` for subscription/OAuth billing
(no behavior change there beyond now-populated token counts) to avoid
implying a dollar cost that doesn't exist for flat-rate billing.

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, no extended
thinking. Root cause was found by comparing real `grok` CLI JSON stream
output (captured directly from a live invocation) against the persisted
`heartbeat_runs.usage_json` row for the same run on a self-hosted
instance, then reading `parse.ts`/`execute.ts` source to confirm the
hardcoded zero values.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change
(`fix/grok-local-usage-cost-tracking`) and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (no
user-facing docs reference this internal usage-reporting behavior)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending at time of writing)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(addressed the one P1 raised — `usageBasis: "per_run"`)
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-11 17:38:23 -07:00
scotttong 815e49bb7c
feat: make chat-style tasks the default experience (#11101) 2026-08-11 09:06:21 -07:00
Nicky Leach 459469638e
feat(adapter-codex-local): add secure device-login building blocks (#11097)
## Thinking Path

> - Paperclip connects AI agents to local and remote runtimes.
> - The Codex local adapter needs a safe device-login flow.
> - A future sandbox integration needs strict prompt validation, secret
protection, cleanup, and private credential storage.
> - This pull request adds tested building blocks for that flow.
> - The result gives a later Daytona integration a clear security
boundary.

## Linked Issues or Issue Description

No public issue covers this change.

**Problem or motivation**

The Codex local adapter has no safe, reusable flow to prove device login
inside an isolated sandbox.

**Proposed solution**

Add parser, runner, credential export, and proof helpers. Validate the
prompt, protect login data, store credentials in a private run-scoped
home, and dispose all sandbox resources.

**Alternatives considered**

Do not connect a production Daytona driver in this change. Use an
injected sandbox driver and focused tests first. This keeps the security
controls testable before live provider integration.

**Roadmap alignment**

The change extends the Codex local adapter. It does not add a core
Paperclip route or duplicate a planned core feature.

**Additional context**

The flow keeps the login URL, code, and token out of logs, results, and
errors. The proof home uses a company-scoped root and a run-scoped
private directory.

## What Changed

- Add a pure parser for the exact Codex device-login URL and one-time
code shape.
- Add a sandbox runner with prompt handling, timeout, cancellation, and
disposal.
- Add a credential export step with company scoping, path checks,
payload checks, private modes, locking, and cleanup.
- Add redacted device-login fixtures and focused tests for parsing,
secret redaction, runner outcomes, credential export, and cleanup.

## Verification

- Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`.
- Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc
--noEmit`.
- Review tests for strict URL and code validation, timeout,
cancellation, disposal, secret redaction, path safety, payload safety,
file modes, and cleanup.

## Risks

- This change provides building blocks, not a live Daytona proof.
- A later integration must connect the runner to a concrete sandbox
driver.
- Credential export depends on existing Codex authentication cache
helpers.
- Incorrect path or payload assumptions can reject valid credentials.

## Model Used

Codex, GPT-5, tool use, code execution, and repository review. The
Paperclip runtime controls the exact context window and reasoning mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-10 16:12:30 -07:00
Nicky Leach f5e9ca3e89
fix(adapters): keep user-scoped env bindings on the agent Test action (#10926)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The agent Test action builds adapter config from the form.
> - The build-config parser kept plain and secret_ref bindings.
> - It dropped user_secret_ref bindings on the create path.
> - This PR shares one parser that keeps every binding shape.
> - The test path now sends the same env binding set that a real run
sees.
> - The benefit is one fix across every adapter build-config path.

## Linked Issues or Issue Description

**What happened?**
The agent Test action dropped a user-scoped env binding in create mode.
The same agent config worked in a real run. Related public PRs: #10115,
#9321, #9921, #8825.

**Expected behavior**
The Test action should keep user-scoped env bindings and resolve them
like a real run.

**Steps to reproduce**
1. Set a user-scoped env binding on an agent config form.
2. Run Test in create mode.
3. The probe runs without the variable.

**Paperclip version or commit**
c09d2509e3

**Deployment mode**
Local dev (pnpm dev)

**Agent adapter(s) involved**
Not adapter-specific (core bug)

**Database mode**
Embedded PGlite (default — DATABASE_URL unset)

**Additional context**
This change is not Claude-specific.

## What Changed

- Added a shared env binding parser in `@paperclipai/adapter-utils`.
- Replaced the eight adapter build-config copies with the shared helper.
- Kept `plain`, `secret_ref`, and `user_secret_ref` bindings intact in
create mode and edit mode.
- Preserved the runtime merge behavior from the earlier env merge
change.

## Verification

- Author-recorded test run:
`packages/adapter-utils/src/env-bindings.test.ts`
- Author-recorded test run:
`packages/adapters/claude-local/src/ui/build-config.test.ts`
- Author-recorded test run: six adapter build-config test files
- Author-recorded typecheck: `tsc --noEmit` for adapter-utils and the
eight adapter packages
- GitHub checks: all required PR checks pass on PR #10926.
- Greptile review: 5/5 with no open comments.

## Risks

- The change touches adapter config assembly.
- A wrong binding shape would change test-time probe input.
- Tests cover the binding types and the create-mode path.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-5, code execution and repo inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes or
confirmed no documentation update is needed
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-05 17:18:41 -07:00
Devin Foley ffd62a4cbb
fix(adapter-utils): carry the workspace origin remote into transported git workspaces (#10873)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent runs on a different host (sandbox or SSH), the adapter
transport copies the local git execution workspace to that host and
syncs changes back after the run
> - The transport materializes the remote copy with `git init` plus a
depth-1 or bundle fetch, so the copy has no `origin` remote and its head
reads as a parentless snapshot commit
> - An agent asked to publish its branch (push it, open a pull request)
sees "no remote, root snapshot" and must hand the publish step back to a
human operator, even when the branch base is a commit the upstream
remote already holds
> - This pull request carries the workspace's `origin` URL
(credential-scrubbed) onto the transported copy as metadata
> - The benefit is that branches produced in transported workspaces stay
publishable by any actor with credentials, while the transport itself
still never fetches or pushes

## Linked Issues or Issue Description

No public issue exists. Description follows the enhancement template:

**What existing behavior does this improve?**

The workspace transport in `@paperclipai/adapter-utils` already copies a
git workspace to the execution host and back. This change improves the
fidelity of that copy: the transported repo keeps the workspace's
`origin` remote instead of losing it.

**Subsystem affected**

Adapter utilities — the sandbox transport
(`withShallowGitWorkspaceClone` in
`packages/adapter-utils/src/git-workspace-sync.ts`) and the SSH
transport (`importGitWorkspaceToSsh` in
`packages/adapter-utils/src/ssh.ts`).

**Current behavior**

The transported copy is built with `git init` plus a depth-1 (sandbox)
or bundle (SSH) fetch. It has no remotes. `git remote -v` is empty and
the head commit reads as a root snapshot with no visible ancestry.
Agents and operators inside the execution host cannot fetch real
ancestry or push a branch, even when the branch base is a commit the
upstream remote already holds.

**Proposed behavior**

The transport reads the source workspace's `origin` URL, scrubs
credentials from it, and configures it on the transported copy. The
sandbox path adds the remote to the fresh clone. The SSH path sets or
adds the remote in the remote setup script, which also covers reused
workspace directories. A workspace with no `origin` transports exactly
as before.

**Reason and benefit**

A branch committed in a transported workspace becomes publishable in
place: the shallow boundary commit already exists on the remote, so a
push pack closes without full local ancestry (a new test locks in this
property). Fetching real ancestry also becomes possible for whoever
holds credentials. Without this, agents must describe their change in a
handoff document and a human must reconstruct the branch by hand.

**Breaking changes**

None. The URL copy is best-effort and metadata-only. The transport never
fetches from or pushes to the remote. The no-remote-git contract holds:
sync-back through the local cwd stays the only cross-run persistence
path, and `packages/adapters/AUTHORING.md` gains a paragraph that makes
the carried-remote nuance explicit.

## What Changed

- `packages/adapter-utils/src/git-workspace-sync.ts`: new
`sanitizeGitRemoteUrl` (strips http(s) userinfo, where tokens can be
embedded; scp-like/ssh forms and filesystem paths pass through) and
`readSanitizedOriginRemoteUrl`; `withShallowGitWorkspaceClone`
configures the scrubbed `origin` on the fresh clone, best-effort.
- `packages/adapter-utils/src/ssh.ts`: `importGitWorkspaceToSsh` sets or
adds the scrubbed `origin` in the remote setup script, non-fatal under
`set -e`.
- `packages/adapter-utils/src/git-workspace-sync.test.ts`: four new
integration cases (remote copied, credentials scrubbed, no-origin
unchanged, push from the shallow clone to an origin that holds the base
commit) plus `sanitizeGitRemoteUrl` unit tests.
- `packages/adapters/AUTHORING.md`: documents that a transported copy
may carry a credential-scrubbed `origin` as metadata, and why this does
not weaken the no-remote-git contract.

## Verification

- `npx vitest run packages/adapter-utils/src/git-workspace-sync.test.ts`
— 12/12 pass (4 new integration cases + sanitizer unit tests).
- `npx vitest run
packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 24/24
pass.
- `npx vitest run packages/adapter-utils/src/ssh-fixture.test.ts` —
16/16 pass, including the `no-remote-git contract` case (a workspace
without `origin` still round-trips with no remote introduced at any
point).
- `node scripts/check-no-git-push.mjs` — passes; this change adds no
push or fetch to adapter/runtime code.
- `pnpm typecheck` in `packages/adapter-utils` — clean.

## Risks

- Low risk. The change is additive metadata on the transported copy
only; failure to record the remote never fails the transport.
- Credential exposure is the real hazard and is handled: http(s)
userinfo is stripped before the URL leaves the host. Non-http forms
(scp-like, `ssh://`) carry no secret in the URL and pass through.
- A reused SSH workspace whose project `origin` changed now gets the
current URL via `set-url` instead of keeping a stale one.

## Model Used

Claude Fable 5 (`claude-fable-5`), Anthropic — extended thinking,
agentic tool use via Claude Code CLI.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-05 10:24:01 -07:00
Joonyoung Park c54936e2e9
fix(openclaw-gateway): use per-agent claimedApiKeyPath in wake text (#4668)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The `openclaw-gateway` adapter wakes a remote agent over a WebSocket
gateway. It sends a wake prompt. That prompt tells the agent which
environment variables to set and which file holds its Paperclip API key.
> - Each agent stores its claimed key in its own JSON file. The adapter
already exposes a `claimedApiKeyPath` config field for this. The field
is documented in `src/index.ts`. It also has an input in the agent
settings UI.
> - `buildWakeText` ignored that field. It hardcoded the shared default
path into the wake prompt text.
> - Every agent therefore read the same key file at wake time. Agents
authenticated as the wrong identity. The first API call failed.
> - This pull request passes `ctx.config.claimedApiKeyPath` into
`buildWakeText`. It uses the existing `resolveClaimedApiKeyPath` helper.
That helper falls back to the documented default.
> - The benefit is that each agent reads its own claimed-key file. Each
agent authenticates as itself.

## Linked Issues or Issue Description

Fixes #10071
Fixes #4976
Fixes #3098
Fixes #8076

These four open issues report the same defect. Earlier duplicates are
already closed: Refs #2561, Refs #2592, Refs #930.

Related pull requests that address the same root problem (duplicate
search):

- #3396 — same core change, no tests
- #3370 — heavier approach, injects `PAPERCLIP_CLAIMED_API_KEY_PATH`
into the wake env and adds server onboarding defaults
- #5970 — renames the config field to `paperclipApiKeyPath`
- #8072 — same core change, bundled with an unrelated protocol-version
change
- #784 — adds shell quoting and preflight instructions
- #3296 — bundled with an unrelated Claude hello-probe fix

## What Changed

- `packages/adapters/openclaw-gateway/src/server/execute.ts`
- `buildWakeText` now accepts `claimedApiKeyPath` as a parameter. It no
longer hardcodes the path.
- The `execute` call site passes
`resolveClaimedApiKeyPath(ctx.config.claimedApiKeyPath)`. That helper
returns the documented default
`~/.openclaw/workspace/paperclip-claimed-api-key.json` when the agent
sets no override.
  - `resolveClaimedApiKeyPath` is now exported so tests can call it.
- `packages/adapters/openclaw-gateway/src/server/execute.test.ts` — adds
`resolveClaimedApiKeyPath` cases: a configured value, an empty string, a
whitespace-only string, `undefined`, `null`, and non-string input.
- `packages/adapters/openclaw-gateway/vitest.config.ts` (new) —
package-level vitest config. It matches the config used by sibling
adapters such as `opencode-local`.
- `vitest.config.ts` (root) — adds the adapter to the workspace project
list.
- `scripts/run-vitest-stable.mjs` — adds
`@paperclipai/adapter-openclaw-gateway` to `nonServerProjects`.
**Maintainer-added during rebase.** The CI test lanes do not run a bare
`vitest`. They call `run-vitest-stable.mjs`, which invokes vitest with
an explicit `--project` allowlist. Without this entry the CI lanes skip
this package, and the root project-list entry alone has no effect on CI.

## Verification

Run the package suite directly:

```
pnpm install --frozen-lockfile
pnpm exec vitest run --project @paperclipai/adapter-openclaw-gateway
```

The suite covers `resolveSessionKey`, `buildAgentParams`, and the new
`resolveClaimedApiKeyPath` cases. The first two already existed in this
file but never executed in CI before this change.

Typecheck the package:

```
pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck
```

Behavioural check, which no automated test covers:

1. Set `claimedApiKeyPath` to a per-agent value such as
`~/.openclaw/workspace/paperclip-keys/<agent>.json` in the agent's
gateway adapter settings.
2. Trigger a wake for that agent.
3. Confirm the rendered wake text names that file. It must not name the
shared default.

Maintainer note: this branch was rebased onto current `master` by a
maintainer. The original branch was two months stale. Only two conflicts
occurred, both additive: the import line and the tail of
`execute.test.ts`, and the project list in the root `vitest.config.ts`.
The `execute.ts` change applied without conflict. CI and Greptile re-run
against the rebased head.

## Risks

- Low for existing deployments. `resolveClaimedApiKeyPath` preserves the
default path exactly. Any agent that never set `claimedApiKeyPath`
receives the same wake text as before.
- The behaviour changes only for agents that already set a per-agent
path. Those agents previously received the wrong instruction. They now
receive the correct one.
- No database, schema, or API surface changes.
- CI now runs this package's test file for the first time. That file
includes the pre-existing `resolveSessionKey` and `buildAgentParams`
tests, which were never executed before.
- Five other adapters (`cursor-cloud`, `cursor-local`, `gemini-local`,
`grok-local`, `pi-local`) sit in the root project list but remain absent
from the CI allowlist. This pull request does not change them. That gap
is tracked separately.

## Model Used

- Contributor's change: Anthropic Claude, model ID `claude-opus-4-7`,
approximately 200K context, extended thinking. Used for triage, patch
authoring, and the original description.
- Rebase, the `run-vitest-stable.mjs` entry, and this description:
Anthropic Claude, model ID `claude-opus-5`, tool use enabled. Run by a
Paperclip maintainer.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
— the branch name carries an internal ticket id. A fork branch cannot be
renamed without opening a new pull request, so this is left as-is. The
internal reference has been removed from the description.
- [ ] I have run tests locally and they pass — the contributor verified
the pre-rebase branch. The rebased head is verified by CI on this pull
request.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes —
`claimedApiKeyPath` is already documented in `src/index.ts` and exposed
in the agent settings UI, so no documentation change is needed
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending the post-rebase run
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending re-review of the rebased head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Pieter (CTO) <pieter@openclaw.local>
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
2026-08-05 00:53:43 -05:00
Nicky Leach 7e592c33fa
feat(codex): add identity-keyed host credential cache (#10853)
## Thinking Path

> - Paperclip manages agent work as issues and pull requests.
> - This change lives in the Codex local adapter path.
> - The host auth flow needs a stable cache per identity.
> - The cache must not change the copy-back path or the default store
overwrite.
> - This pull request adds that cache and keeps the existing flow
intact.
> - The benefit is repeatable host auth with safer identity scoping.

## Linked Issues or Issue Description

**Subsystem affected**
- packages/adapters: Codex local adapter and host auth flow

**Problem or motivation**
- The host auth flow needs one usable credential per identity.
- The current flow does not keep that identity state in a separate
cache.

**Proposed solution**
- Add an identity-keyed host credential cache.
- Keep the cache company scoped.
- Add an opt-in seed mode for the merge decision helper.
- Write the cache at copy-back time without changing the default store
overwrite.

**Alternatives considered**
- Store the cache in the instance-global root.
- Seed the host default store through an environment flag.
- Both choices weaken isolation or caller control, so I did not use
them.

**Roadmap alignment**
- I checked `ROADMAP.md`.
- I found no direct overlap with an active roadmap item.
- The change fits the core auth and secrets direction.

**Additional context**
- Local tests passed before I opened this pull request.
- I found no direct duplicate pull request for this branch.

## What Changed

- Added `codex-auth-cache.ts` for company scoped cache storage and
identity anchored vending.
- Added `codex-auth-merge-decision.cjs` support for a leading
`--seed-if-dest-absent` flag.
- Updated `codex-auth-copyback.ts` to write the cache at teardown under
the merge lock.
- Wired `execute.ts` to use the identity anchored vend before the
managed home seed step.
- Added `CODEX-AUTH-CACHE.md` for the cache rules, directions, state
matrix, off switch, and clear action.
- Added and updated tests for cache behavior, merge decision flow, and
copy-back flow.

## Verification

- `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`.
- `node_modules/.bin/vitest run --project @paperclipai/adapter-utils
workspace-restore-merge`.
- `tsc --noEmit` in `packages/adapters/codex-local`.
- `git log --oneline
origin/master..origin/feat/codex-auth-identity-cache` shows the expected
commits.

## Risks

- This change touches credential storage.
- A mistake in the cache path could break host auth for one identity.
- The tests reduce this risk, but the surface stays sensitive.

## Model Used

- OpenAI Codex, GPT-5, tool use enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used with version and capability
details
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues with `Fixes:` / `Closes:` /
`Refs:` or described the issue in the pull request body
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-04 16:33:54 -07:00
Devin Foley f0b06d2de9
feat(claude): environment-aware test-environment probe and claude-local CI coverage (#10833)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The claude_local adapter runs Claude Code on sandbox execution
targets, and operators verify an agent's configuration with the
test-environment probe before running it
> - Real runs merge the selected environment's env vars (secret refs
included) under the agent's adapter config env, but the probe built its
config from the adapter config alone — so environment-level auth worked
in runs while the Test button reported missing auth, and a dropped
secret binding passed silently
> - The claude env-test hints also did not recognize
`CLAUDE_CODE_OAUTH_TOKEN` even though the CLI accepts it, and a hello
probe that hit the subscription usage limit reported a hard failure
although authentication worked
> - Separately, the claude-local package test suites were absent from
the CI project list, so two suites drifted broken without notice
> - This pull request makes the probe resolve the same layered env as a
real run, adds the missing auth hint, classifies usage-limit probe
results as a warning, repairs the drifted suites, and turns the
claude-local project on in CI
> - The benefit is a Test button that tells the truth about
environment-level configuration, and a test suite that actually gates
the claude-local adapter

## Linked Issues or Issue Description

No public issue exists; related open PRs: Refs #9488 (recognizes
CLAUDE_CODE_OAUTH_TOKEN in environment checks — overlaps with the
auth-hint portion of this PR via a differently named check; it does not
cover the environment-envVars probe merge, the usage-limit
classification, or the CI coverage), Refs #9933 (live credential
validation in environment checks — complementary, no file-level conflict
with the route change).

The underlying problem, following the enhancement template:

**Current behavior**

The test-environment route builds the probe config from the agent's
adapterConfig only. Real runs merge the selected environment's envVars
under the agent env, so environment-level env vars (including auth such
as `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` bound as environment
secrets) work in runs while "Test environment" cannot see them, and a
missing secret binding passes silently. The claude env-test hints do not
recognize `CLAUDE_CODE_OAUTH_TOKEN`. A hello probe that hits the
subscription usage limit reports a hard `claude_hello_probe_failed`. The
claude-local package test suites do not run in CI, and two of them are
stale.

**Proposed behavior**

The probe resolves the selected environment's envVars
(environment-consumer secret bindings included) and merges them under
the agent config env with the run-path precedence; missing bindings
surface as an explicit error check that fails the test. The env-test
emits a `claude_oauth_token_configured` info check when that variable is
set. Usage-limit probe results classify as a
`claude_hello_probe_usage_limited` warning because auth works and only
the usage window is spent. The claude-local suites run in CI. Docs state
the resulting facts.

**Reason and benefit**

The Test button should tell the truth: it previously contradicted run
behavior for environment-level configuration and hid broken secret
bindings. Enabling the package suites in CI prevents further silent
drift — two suites were already broken on master without anyone
noticing.

**Breaking changes**

None. Runs are unchanged. The probe route only adds env layers and
checks; setups without environment envVars behave exactly as before.

## What Changed

- `server/src/routes/agents.ts`: the test-environment route resolves the
selected environment's envVars (forbidden keys stripped,
environment-consumer secret context) and merges them under the agent
adapterConfig env, mirroring `resolveExecutionRunAdapterConfig`
precedence. Missing secret bindings are skipped, reported as an
`environment_env_binding_missing` error check, and fail the test —
matching the `ConfigurationIncompleteFailure` a real dispatch would
raise.
- `packages/adapters/claude-local/src/server/test.ts`: new
`claude_oauth_token_configured` info hint between the API-key warning
and the subscription fallback; hello-probe classification gains a
`claude_hello_probe_usage_limited` warning for provider-quota results
(previously a hard `claude_hello_probe_failed`).
- `scripts/run-vitest-stable.mjs`: add
`@paperclipai/adapter-claude-local` to `nonServerProjects` so CI runs
the package suites.
- `packages/adapters/claude-local/src/server/execute.remote.test.ts`:
assert both runtime asset syncs (skills and mcp-config); the suite
predated the mcp-config asset.
- `packages/adapters/claude-local/src/server/test.probe.test.ts`:
usage-limit fixture now expects the usage-limited warning; new fixture
covers the genuine transient path (529 overloaded); new tests cover the
token hint and API-key precedence.
- `server/src/__tests__/agent-test-environment-routes.test.ts`: new
tests for the env merge (agent wins on conflict, forbidden key
filtered), missing-binding reporting, and the no-execution-target
fallback path.
- `docs/adapters/claude-local.md`, `docs/adapters/overview.md`: state
the auth-input facts (API key or oauth token wins over stored logins;
snapshot-owns-auth applies when neither is configured) and describe the
environment-aware Test behavior.

## Verification

- `npx vitest run --project @paperclipai/adapter-claude-local` — 131
tests pass (both drifted suites repaired; they fail on master today).
- `npx vitest run
server/src/__tests__/agent-test-environment-routes.test.ts` — 7 tests
pass.
- `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` —
passes with the added project.
- `pnpm typecheck` in `server/` and `packages/adapters/claude-local/` —
clean.

## Risks

- Low risk. The run path is untouched; the probe route change is
additive and inert when the environment has no envVars.
- The probe now performs environment-consumer secret resolution at test
time; access is authorized per binding exactly as at run time, and the
audit consumer is the environment (as before for adapter-config
resolution).
- Enabling the claude-local project in CI adds about 2 seconds of vitest
wall time to the general workspaces group and could surface future
regressions in that package — which is the point.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic
tool use via Claude Code (CLI).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-04 10:31:14 -07:00
Devin Foley bd86dbe41b
fix(codex-local): detect server-visible Codex credentials in ACP environment test (#10703)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The codex-local adapter runs Codex agents through the ACP lane and
offers a "Test" button on the agent configuration page to validate the
environment
> - The environment test used its own ad-hoc credential probe, while
real dispatch uses the shared `evaluateCodexCredentialReadiness`
predicate in `codex-home.ts`
> - The two paths disagreed: a user with valid Codex subscription auth
in the shared, server-visible Codex home still saw "No Codex ACP
credentials were detected"
> - The old warning also suggested `codex login` without explaining that
a `/login` in a separate Codex or chat session does not authenticate the
Paperclip server process
> - This pull request makes `testCodexAcpEnvironment` use the same
shared readiness predicate as real dispatch and rewords the warning to
name the server credential boundary
> - The benefit is that the Test button now agrees with what dispatch
will actually do, and the warning tells the user exactly which process
needs the credentials

## Linked Issues or Issue Description

No public GitHub issue exists for this bug. Related credential-handling
work, not duplicates: Refs #10160 (classifies OpenAI invalid-key 401s at
dispatch time) and Refs #9598 (classifies Codex refresh auth failures).
Both cover dispatch-time failures; this PR fixes the pre-dispatch
environment test.

Description follows the bug report template:

**What happened?**

A user configured a Codex ACP agent and had already authenticated Codex
(subscription auth present in the shared Codex home visible to the
Paperclip server). Clicking "Test" on the agent configuration page still
reported:

`warn: No Codex ACP credentials were detected. Hint: Set OPENAI_API_KEY
or run codex login before starting a Codex ACP agent.`

**Expected behavior**

The environment test should detect the same credentials that real agent
dispatch would use. When shared managed Codex auth is available, the
test should pass with an informational check instead of warning. When
credentials really are missing, the warning should explain that the
Paperclip server process is the one that needs them.

**Steps to reproduce**

1. Run the Paperclip server as an OS user whose shared Codex home
contains valid subscription `auth.json` (no `OPENAI_API_KEY` in the
adapter env or server env).
2. Configure an agent with the codex-local adapter using the ACP engine.
3. Click "Test" on the agent configuration page.
4. Observe the `codex_acp_credentials_missing` warning even though
dispatch would succeed.

**Agent adapter(s) involved**

Codex

**Additional context**

The confusion was amplified by the hint: users had run `/login` in a
Codex chat session and assumed the server was authenticated. That login
lives in a different process and home directory, so the server never saw
it.

## What Changed

- `testCodexAcpEnvironment`
(packages/adapters/codex-local/src/server/acp.ts) now calls the shared
`evaluateCodexCredentialReadiness` predicate from `codex-home.ts`
instead of a local ad-hoc `hasCodexNativeCredentials` probe, so the Test
button and real dispatch agree.
- An explicit empty `OPENAI_API_KEY` in the adapter config env no longer
falls through to the server environment key.
- An externally managed `CODEX_HOME` override is now reported as its own
informational check (`codex_acp_external_home_configured`).
- The `codex_acp_credentials_missing` warning now says the credentials
must be visible to the Paperclip server, and the hint explains that a
`/login` in a separate Codex or chat session does not authenticate the
server.
- Removed the now-unused `hasCodexNativeCredentials` helper.
- Added two regression tests: shared managed Codex auth is detected (no
false warning), and the missing-credentials warning carries the new
server-boundary wording.

## Verification

- `pnpm --filter @paperclip/adapter-codex-local test --
src/server/acp.test.ts` — the two new tests cover the shared-home
detection branch and the new warning wording; the existing ACP lane
tests cover the API-key and remote-target branches.
- Manual: with subscription auth in the server-visible shared Codex home
and no `OPENAI_API_KEY`, the agent configuration Test now reports
`codex_acp_native_auth_detected` (info) instead of
`codex_acp_credentials_missing` (warn).

## Risks

- Low risk. The change only affects the environment test path, not
dispatch. The readiness predicate is the same one dispatch already uses,
so drift between the two paths is now structurally prevented.
- Behavioral shift: an explicit empty adapter `OPENAI_API_KEY` no longer
silently falls back to the server env key in the test result. This
matches dispatch behavior and is intentional.

## Model Used

- Implementation authored by OpenAI Codex (gpt-5.5) running through the
Codex ACP lane with tool use.
- PR preparation, rebase onto master, and review fix-up by Anthropic
Claude (Claude Code CLI agent, extended thinking, tool use).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Cody <noreply@paperclip.ing>
2026-08-02 16:55:53 -07:00
Devin Foley ee9d907d01
fix(codex): do not inject a duplicate --skip-git-repo-check for sandbox runs (#10595)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A Codex agent runs `codex exec`, and the adapter assembles its
argument vector from the agent's config plus execution-context options
> - For sandbox execution the adapter injects `--skip-git-repo-check`,
because a headless remote workspace has no git trust prompt to answer
> - The adapter also appends the operator's `extraArgs` verbatim, so an
agent that already lists `--skip-git-repo-check` in its config gets the
flag twice on a sandbox run
> - `codex exec` rejects a repeated `--skip-git-repo-check` and exits
with code 2, which the adapter surfaces as `adapter_failed` before any
work runs
> - This pull request skips the sandbox injection when the operator's
args already carry the flag
> - The benefit is that a common, harmless-looking config no longer
crashes every sandbox run

## Linked Issues or Issue Description

**What happened?**

A `codex_local` agent configured with `extraArgs:
["--skip-git-repo-check"]` fails on every sandbox run:

```
error: the argument '--skip-git-repo-check' cannot be used multiple times

Usage: codex exec [OPTIONS] [PROMPT]
```

The adapter reports `stopReason: "adapter_failed"` (Codex exited with
code 2). The flag appears twice in the argv: once injected by the
adapter for sandbox execution, once from the operator's `extraArgs`.

**Steps to reproduce**

1. Configure a `codex_local` agent with `extraArgs:
["--skip-git-repo-check"]` (or the legacy `args` field).
2. Point it at a sandbox environment.
3. Start a run — `codex exec` aborts immediately on the duplicate flag.

**Expected behavior**

The run launches with a single `--skip-git-repo-check`. An operator
listing the flag the adapter already injects should be a no-op, not a
hard failure.

**Paperclip version**

Current `master`.

**Deployment mode**

Any deployment running Codex agents in sandbox environments.

## What Changed

- `buildCodexExecArgs` no longer pushes the sandbox
`--skip-git-repo-check` when the resolved args (`extraArgs`, or the
legacy `args` fallback) already contain it. The operator's copy stands;
the argv carries the flag exactly once. Non-sandbox runs and configs
without the flag are unchanged.

## Verification

- `cd packages/adapters/codex-local && pnpm vitest run
src/server/codex-args.test.ts` — new cases: `extraArgs` already carrying
the flag (single occurrence), the legacy `args` field carrying it
(single occurrence), and the operator's flag preserved when the sandbox
injection is not requested. Existing "adds --skip-git-repo-check when
requested" case unchanged.
- `cd packages/adapters/codex-local && pnpm vitest run` — full package
suite (218 tests).
- `pnpm run typecheck` in the package.

## Risks

- Low. The change only suppresses a duplicate of a single, idempotent
flag; it never removes an operator-supplied argument and never adds one
that was not already going to be present.

## Model Used

Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended
thinking, agentic tool use (file edits, vitest/tsc runs). No other
models involved.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-31 20:37:51 -07:00
Devin Foley c0b875c46c
fix(codex): let sandbox runs use the sandbox image's own Codex login (#10582)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Codex agents can run inside sandbox environments, and operators can
bake a Codex login into the sandbox image during interactive image setup
> - Two credential gates (the control plane's pre-dispatch
configuration-incomplete gate and the adapter's execute-time fail-fast)
required host-side Codex credentials — a usable `auth.json` in the
managed home or a configured `OPENAI_API_KEY` — regardless of where the
run executes
> - On managed cloud hosts a local Codex login never exists, so every
sandbox run of a Codex agent failed immediately with "configuration
incomplete: no Codex credentials available for managed home …", even
though the adapter's inbound auth merge already supports the image-login
case end to end
> - This pull request makes the execute-time gate probe the sandbox for
its own `~/.codex/auth.json` before failing, and exempts
sandbox-destined runs from the pre-dispatch host check
> - The benefit is that a sandbox image signed in to Codex is a
first-class credential source, matching what the auth-merge,
precedence-warning, and copy-back machinery were already built for

## Linked Issues or Issue Description

**What happened?**

Running a `codex_local` agent in a sandbox environment whose image
carries a Codex login failed instantly with `configuration incomplete:
no Codex credentials available for managed home "…/codex-home". Sign in
to Codex on the host with a ChatGPT subscription, or bind a per-agent
OPENAI_API_KEY secret for this agent.` The host has no Codex login and
never will on a managed cloud deployment; the sandbox's own login was
never consulted.

**Steps to reproduce**

1. Configure a sandbox environment and capture a custom image after
signing in to Codex inside the interactive image setup.
2. Create a `codex_local` agent that uses that environment, on a host
with no Codex login and no `OPENAI_API_KEY` bound.
3. Start a run: it fails pre-dispatch with the configuration-incomplete
blocker above.

**Expected behavior**

The run launches and Codex authenticates with the sandbox image's own
login, the same way the adapter's host↔sandbox auth merge already keeps
the sandbox credential when the host ships none. A run should only fail
fast when neither the host, a bound `OPENAI_API_KEY`, nor the sandbox
has credentials.

**Paperclip version**

Current `master` (cloud image deployments).

**Deployment mode**

Managed cloud stacks (any deployment where the server host has no local
Codex login).

## What Changed

- Extracted the adapter's execute-time gate into
`assertCodexCredentialsLaunchable`: when host readiness fails and the
target is a sandbox, it probes `~/.codex/auth.json` in the sandbox (same
command the auth-precedence warning uses) and proceeds with a log line
naming the credential source; when the sandbox has no login either, the
error now names all three remediation options (sandbox image sign-in,
per-agent `OPENAI_API_KEY`, host sign-in). Non-sandbox targets keep
today's strict behavior byte-for-byte.
- The control plane's pre-dispatch gate in
`resolveExecutionRunAdapterConfig` now takes the selected environment's
driver and skips the host-credential check for sandbox-destined runs —
only the adapter can probe the sandbox once it is up, so the
execute-time gate is the authority there. Non-sandbox runs keep the
early, well-attributed configuration-incomplete blocker.
- The codex Test flow needed no change: it already seeds host
credentials only when they exist and otherwise leaves the sandbox's
`CODEX_HOME` alone; this aligns the run path with it.

## Verification

- `cd packages/adapters/codex-local && pnpm vitest run` — 210 tests,
including new gate cases: sandbox login present (proceeds + logs
source), sandbox and host both credential-less (fails with the extended
message), non-sandbox target (strict host requirement kept, no sandbox
probe), per-agent API key (no probe at all).
- `cd server && pnpm vitest run
src/__tests__/heartbeat-project-env.test.ts
src/__tests__/codex-local-adapter-environment.test.ts` — includes the
new sandbox-exemption case next to the existing blocker tests.
- `pnpm run typecheck` in `server` and `packages/adapters/codex-local`.

## Risks

- Sandbox-destined misconfigurations (no credentials anywhere) now
surface at adapter execute time instead of pre-dispatch, so they read as
an adapter failure with a precise message rather than a
configuration-incomplete blocker. The trade-off is deliberate: the
sandbox must be up to know whether credentials exist, and the failure
message names the exact remediations.
- The sandbox probe adds one short (5s-capped) shell command to sandbox
runs whose host has no credentials; runs with host credentials or a
bound key are untouched.
- Self-hosted behavior is unchanged for local and SSH targets.

## Model Used

Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended
thinking, agentic tool use (file edits, vitest/tsc runs). No other
models involved.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-31 18:36:36 -07:00
Michael Nguyen d295251550
feat(adapter-claude): add Claude Sonnet 5 to the static model fallback (#10280)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents pick their model from a dropdown in agent config, populated
per-adapter by `listAdapterModels()` → each adapter's live provider
catalog merged over a static fallback list
> - For `claude_local`, newer model ids only reach the dropdown via the
live Anthropic `/v1/models` fetch, which needs a server
`ANTHROPIC_API_KEY`, a <5s round-trip, non-Bedrock mode, and account
entitlement; on any miss it silently falls back to the static `models`
array
> - Claude Sonnet 5 (`claude-sonnet-5`) is a current flagship but was
absent from that static fallback, so it appeared only when live
discovery happened to succeed — i.e. "the newest model doesn't
consistently show up"
> - This pull request adds `claude-sonnet-5` to the `claude_local`
static model list so it is selectable regardless of the live-discovery
path
> - The benefit is a consistent, reliable dropdown that no longer
depends on a flaky live fetch to surface a shipped flagship model

## Linked Issues or Issue Description

No public GitHub issue. The bug is described inline following the
bug-report template:

**What happened**
The `claude_local` agent-config model dropdown intermittently omitted
Claude Sonnet 5. `claude-sonnet-5` was missing from the adapter's static
fallback `models` array (`packages/adapters/claude-local/src/index.ts`),
so it only surfaced when the live Anthropic `/v1/models` discovery
happened to succeed.

**Expected behavior**
Claude Sonnet 5 is a shipped flagship model and should always be
selectable in the dropdown, independent of whether live discovery
succeeds.

**Steps to reproduce**
1. Run the server without a working live Anthropic `/v1/models` path (no
`ANTHROPIC_API_KEY`, Bedrock mode, a discovery timeout, or a cache
miss).
2. Open agent config for a `claude_local` agent and inspect the model
dropdown.
3. Observe that `claude-sonnet-5` is absent because the static fallback
list omitted it.

**Deployment mode**
Self-hosted / local adapter (`claude_local`); the server process reads
`ANTHROPIC_API_KEY` from its environment.

## What Changed

- Added `{ id: "claude-sonnet-5", label: "Claude Sonnet 5" }` to the
`claude_local` static `models` fallback, immediately after
`claude-opus-4-8` (so Opus 4.8 stays the default first option).
- Added an explicit regression assertion in
`server/src/__tests__/adapter-models.test.ts` that `claude-sonnet-5` is
present in the `claude_local` fallback when live discovery is
unavailable.

## Verification

- `pnpm -C server exec vitest run src/__tests__/adapter-models.test.ts
-t "claude fallback"` — **passes** (the new `claude-sonnet-5` assertion
included).
- Reviewed the consuming tests: the fallback test also asserts
`models[0]?.id === "claude-opus-4-8"` (still index 0 — Sonnet 5 is index
1, unaffected); `adapter-registry.test.ts` reads `builtIn?.models`
dynamically, so no exact-array snapshot breaks.
- Change is a single static-data addition plus a test assertion; no
control-flow change.

## Risks

- Low risk. Pure additive change to a fallback list; no control-flow
change. Worst case is an id that a given account isn't entitled to,
which the existing "current"/manual-model UI paths already tolerate.

## Model Used

Claude (Anthropic), model id `claude-opus-4-8` (Opus 4.8), extended
thinking + tool use, run as the Paperclip CTO agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (branch is the assigned
execution-workspace branch and cannot be renamed this run)
- [x] I have run tests locally and they pass (server adapter-models
"claude fallback" case)
- [x] I have added or updated tests where applicable (explicit
`claude-sonnet-5` fallback assertion)
- [x] I have updated relevant documentation to reflect my changes (n/a —
no docs reference this list)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending CI)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending review)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-31 13:15:42 -05:00
Noah Kellner bd53b99686
feat(gemini-local): export detectGeminiQuotaExhausted for plugin adapter reuse (#3416)
## Thinking Path

> - Paperclip orchestrates AI agents via pluggable server-side adapters,
one per provider / CLI backend
> - Each adapter package exposes a server entry point
(`@paperclipai/adapter-<name>/server`) that downstream consumers —
including plugin adapters that wrap or extend the stock behavior —
re-use for helper functions
> - `gemini-local`'s server entry re-exports a curated set of parse
helpers from `./parse.js` (`parseGeminiJsonl`,
`isGeminiUnknownSessionError`, `describeGeminiFailure`,
`detectGeminiAuthRequired`, `isGeminiTurnLimitResult`) so consumers can
classify provider output without reaching into the package's internals
> - `detectGeminiQuotaExhausted` lives in the same `parse.ts` file next
to `detectGeminiAuthRequired`, is defined with `export function`, and is
the package's single source of truth for "is this output a Gemini quota
hit?" — but it is missing from the entry-point re-export block
> - As a result, any consumer that wants to classify quota exhaustion
either has to deep-import from `./server/parse.js` (brittle against
future `exports`-field changes) or reimplement the regex locally (drift
risk against the authoritative heuristic)
> - This pull request adds `detectGeminiQuotaExhausted` to the existing
re-export block, placed next to its thematic sibling
`detectGeminiAuthRequired`, with no other changes
> - The benefit is one extra supported public symbol on
`@paperclipai/adapter-gemini-local/server` — a purely additive ergonomic
improvement with no behavior change and no existing consumer impact

## What Changed

- `packages/adapters/gemini-local/src/server/index.ts`: added
`detectGeminiQuotaExhausted` to the `export { ... } from "./parse.js"`
block, inserted between `detectGeminiAuthRequired` and
`isGeminiTurnLimitResult` (thematic grouping — both `detect*` helpers)

## Verification

Local verification against the branch commit (base: `upstream/master` at
`b649bd45`):

```
$ pnpm --filter @paperclipai/adapter-gemini-local typecheck
> @paperclipai/adapter-gemini-local@0.3.1 typecheck
> tsc --noEmit
(exit 0)

$ pnpm --filter @paperclipai/adapter-gemini-local build
> @paperclipai/adapter-gemini-local@0.3.1 build
> tsc
(exit 0)

$ cd server && pnpm exec vitest run src/__tests__/gemini-local-execute.test.ts
 RUN  v3.2.4

 ✓ src/__tests__/gemini-local-execute.test.ts (3 tests) 1487ms
   ✓ gemini execute > passes prompt via --prompt and injects paperclip env vars
   ✓ gemini execute > always passes --approval-mode yolo
   ✓ gemini execute > uses a compact wake delta instead of the full heartbeat prompt when resuming a session

 Test Files  1 passed (1)
      Tests  3 passed (3)
```

Existing-consumer check — all references to `detectGeminiQuotaExhausted`
anywhere in the tree:

```
packages/adapters/gemini-local/src/server/index.ts:9    (this PR's new re-export)
packages/adapters/gemini-local/src/server/parse.ts:253  (the definition)
packages/adapters/gemini-local/src/server/test.ts:19    (intra-package import from "./parse.js")
packages/adapters/gemini-local/src/server/test.ts:174   (intra-package usage)
```

No cross-package consumer references the symbol today, so the new
re-export cannot break any existing import. It is strictly additive to
the public surface of `@paperclipai/adapter-gemini-local/server`.

## Risks

None. Purely additive re-export of a symbol that is already a named
export on `./parse.ts`. The clean-success path, failure paths, and all
other adapter behavior are untouched. No existing consumer is affected.

## Model Used

- **Provider**: Anthropic
- **Model**: Claude Opus 4.6 (1M context)
- **Interface**: Claude Code CLI
- **Role**: Upstream state verification (grep + diff against current
master), PR drafting against the `CONTRIBUTING.md` template, local
typecheck / build / test execution
- **Reasoning Mode**: Extended thinking enabled
- **Human oversight**: Noah Kellner reviewed the one-line re-export
addition and approved submission

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Noah Kellner <noah.kellner@xenoscloud.com>
2026-07-31 13:00:26 -05:00
Dotta 170c1e5adb
fix(claude-local): trust Paperclip URLs in allowlist sandbox (#10438)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Local adapters can confine agent processes with filesystem and
network sandbox policies
> - Allowlist confinement must still let an agent reach Paperclip's own
control plane and managed MCP servers
> - The Codex local adapter already marks those runtime-owned endpoints
as trusted, but the Claude local adapter did not
> - As a result, allowlisted Claude runs could not call their own
Paperclip API unless operators duplicated runtime URLs manually
> - This pull request mirrors the trusted-URL wiring in the Claude local
adapter and adds a regression test at the process-execution boundary
> - The benefit is that confined Claude agents retain their required
control-plane access without broadening the operator-managed network
allowlist

## Linked Issues or Issue Description

Refs #447

No exact public issue was found. The underlying bug is:

- **Observed:** with `claude_local` configured for local-process
`networkScope: "allowlist"`, the sandbox options omitted the
runtime-owned Paperclip API and MCP server URLs. Requests to those
endpoints could therefore be denied by confinement.
- **Expected:** Paperclip's own API URL and managed MCP server URLs are
passed as trusted sandbox targets, matching `codex_local` behavior.
- **Steps to reproduce:** configure a local Claude agent with allowlist
network confinement, omit the runtime Paperclip API URL from the
operator allowlist, and have the agent call its injected
`PAPERCLIP_API_URL`.
- **Version:** `ca92f727c5` (`origin/master` at branch creation).
- **Deployment mode:** local-process sandbox confinement in a local
deployment.

## What Changed

- Added the injected Paperclip API URL and runtime MCP server URLs to
`claude_local` sandbox `networkTrustedUrls`, filtering empty values.
- Added a unit test proving allowlist sandbox construction trusts
`PAPERCLIP_API_URL`.

## Verification

- `pnpm exec vitest run
packages/adapters/claude-local/src/server/execute.acp-fallback.test.ts`
- `pnpm --filter @paperclipai/adapter-claude-local typecheck`

## Risks

- Low risk. The change only affects confined local Claude processes and
trusts runtime-owned endpoints already injected by Paperclip.
- Operator-configured `networkAllowlist` behavior is unchanged; the
added URLs use the sandbox's separate trusted-target path.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex based on GPT-5 (exact serving snapshot and context-window
size are not exposed), with reasoning, repository inspection, shell tool
use, code editing, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (not
applicable: behavior parity fix with no user-facing configuration
change)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-29 13:23:09 -05:00
dependabot[bot] 5b34d265a4
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.59.0 to 0.63.0 (#10302)
Bumps
[@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp)
from 0.59.0 to 0.63.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@​agentclientprotocol/claude-agent-acp's
releases</a>.</em></p>
<blockquote>
<h2>v0.63.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.62.0...v0.63.0">0.63.0</a>
(2026-07-27)</h2>
<h3>Features</h3>
<ul>
<li>Update to claude agent sdk v0.3.220 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/921">#921</a>)
(<a
href="4c7b897183">4c7b897</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Only resolve a denied tool call the client was told about (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/923">#923</a>)
(<a
href="8f67b6a92b">8f67b6a</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/918">#918</a></li>
<li>Report tool_progress heartbeats against the tool call they describe
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/916">#916</a>)
(<a
href="5559ba890c">5559ba8</a>)</li>
<li><strong>tools:</strong> key Bash terminal metas off the announced
tool_use id (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/917">#917</a>)
(<a
href="d0604140f9">d060414</a>)</li>
</ul>
<h2>v0.62.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.61.0...v0.62.0">0.62.0</a>
(2026-07-24)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump <code>@​hono/node-server</code> from
1.19.14 to 1.19.15 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/908">#908</a>)
(<a
href="51936049e1">5193604</a>)</li>
<li><strong>deps:</strong> Bump media-typer from 1.1.0 to 1.1.1 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/909">#909</a>)
(<a
href="5d35001563">5d35001</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 2 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/900">#900</a>)
(<a
href="809d41c6b7">809d41c</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 2 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/907">#907</a>)
(<a
href="14d06273c0">14d0627</a>)</li>
<li>Update to claude-agent-sdk 0.3.218 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/904">#904</a>)
(<a
href="8cbaf97254">8cbaf97</a>)</li>
</ul>
<h2>v0.61.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.60.0...v0.61.0">0.61.0</a>
(2026-07-22)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump actions/setup-node from 6.4.0 to 7.0.0
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/897">#897</a>)
(<a
href="d9bd36d8b0">d9bd36d</a>)</li>
<li><strong>deps:</strong> Update to
<code>@​anthropic-ai/claude-agent-sdk</code> 0.3.217 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/899">#899</a>)
(<a
href="edf3af043b">edf3af0</a>)</li>
</ul>
<h2>v0.60.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.59.0...v0.60.0">0.60.0</a>
(2026-07-20)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Update to claude-agent-sdk 0.3.215 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/890">#890</a>)
(<a
href="92548f0435">92548f0</a>)</li>
<li>implement configurable LLM providers (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/853">#853</a>)
(<a
href="82cd692e50">82cd692</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>parse Agent/Task trailers without regex (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/879">#879</a>)
(<a
href="06c3d7bdbd">06c3d7b</a>)</li>
<li>remove ~15s stall on session/new and model switch by seeding the
context window synchronously (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/894">#894</a>)
(<a
href="ff9b96d462">ff9b96d</a>)</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@​agentclientprotocol/claude-agent-acp's
changelog</a>.</em></p>
<blockquote>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.62.0...v0.63.0">0.63.0</a>
(2026-07-27)</h2>
<h3>Features</h3>
<ul>
<li>Update to claude agent sdk v0.3.220 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/921">#921</a>)
(<a
href="4c7b897183">4c7b897</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Only resolve a denied tool call the client was told about (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/923">#923</a>)
(<a
href="8f67b6a92b">8f67b6a</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/918">#918</a></li>
<li>Report tool_progress heartbeats against the tool call they describe
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/916">#916</a>)
(<a
href="5559ba890c">5559ba8</a>)</li>
<li><strong>tools:</strong> key Bash terminal metas off the announced
tool_use id (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/917">#917</a>)
(<a
href="d0604140f9">d060414</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.61.0...v0.62.0">0.62.0</a>
(2026-07-24)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump <code>@​hono/node-server</code> from
1.19.14 to 1.19.15 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/908">#908</a>)
(<a
href="51936049e1">5193604</a>)</li>
<li><strong>deps:</strong> Bump media-typer from 1.1.0 to 1.1.1 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/909">#909</a>)
(<a
href="5d35001563">5d35001</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 2 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/900">#900</a>)
(<a
href="809d41c6b7">809d41c</a>)</li>
<li><strong>deps:</strong> Bump the minor group with 2 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/907">#907</a>)
(<a
href="14d06273c0">14d0627</a>)</li>
<li>Update to claude-agent-sdk 0.3.218 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/904">#904</a>)
(<a
href="8cbaf97254">8cbaf97</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.60.0...v0.61.0">0.61.0</a>
(2026-07-22)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Bump actions/setup-node from 6.4.0 to 7.0.0
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/897">#897</a>)
(<a
href="d9bd36d8b0">d9bd36d</a>)</li>
<li><strong>deps:</strong> Update to
<code>@​anthropic-ai/claude-agent-sdk</code> 0.3.217 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/899">#899</a>)
(<a
href="edf3af043b">edf3af0</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.59.0...v0.60.0">0.60.0</a>
(2026-07-20)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> Update to claude-agent-sdk 0.3.215 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/890">#890</a>)
(<a
href="92548f0435">92548f0</a>)</li>
<li>implement configurable LLM providers (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/853">#853</a>)
(<a
href="82cd692e50">82cd692</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>parse Agent/Task trailers without regex (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/879">#879</a>)
(<a
href="06c3d7bdbd">06c3d7b</a>)</li>
<li>remove ~15s stall on session/new and model switch by seeding the
context window synchronously (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/894">#894</a>)
(<a
href="ff9b96d462">ff9b96d</a>)</li>
<li>Silence missing PostToolUse callbacks (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/895">#895</a>)
(<a
href="1887ada215">1887ada</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/889">#889</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="15979bba79"><code>15979bb</code></a>
chore(main): release 0.63.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/922">#922</a>)</li>
<li><a
href="8f67b6a92b"><code>8f67b6a</code></a>
fix: Only resolve a denied tool call the client was told about (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/923">#923</a>)</li>
<li><a
href="5559ba890c"><code>5559ba8</code></a>
fix: Report tool_progress heartbeats against the tool call they describe
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/916">#916</a>)</li>
<li><a
href="d0604140f9"><code>d060414</code></a>
fix(tools): key Bash terminal metas off the announced tool_use id (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/917">#917</a>)</li>
<li><a
href="4c7b897183"><code>4c7b897</code></a>
feat: Update to claude agent sdk v0.3.220 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/921">#921</a>)</li>
<li><a
href="8663170f18"><code>8663170</code></a>
Add structured Bash titles and nested subagent transcripts (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/881">#881</a>)</li>
<li><a
href="53a0c36ce3"><code>53a0c36</code></a>
chore(main): release 0.62.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/901">#901</a>)</li>
<li><a
href="c91f943f86"><code>c91f943</code></a>
<code>@​anthropic-ai/claude-agent-sdk</code> 0.3.218 -&gt; 0.3.219 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/912">#912</a>)</li>
<li><a
href="5d35001563"><code>5d35001</code></a>
feat(deps): Bump media-typer from 1.1.0 to 1.1.1 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/909">#909</a>)</li>
<li><a
href="51936049e1"><code>5193604</code></a>
feat(deps): Bump <code>@​hono/node-server</code> from 1.19.14 to 1.19.15
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/908">#908</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.59.0...v0.63.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-29 09:33:08 -07:00
Michael Nguyen 4eace88f6b
feat(adapter-claude): add Claude Opus 5 to the static model fallback (#10327)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents pick their model from a dropdown in agent config, populated
per-adapter by `listAdapterModels()` → each adapter's live provider
catalog merged over a static fallback list
> - For `claude_local`, newer model ids only reach the dropdown via the
live Anthropic `/v1/models` fetch, which needs a server
`ANTHROPIC_API_KEY`, a <5s round-trip, non-Bedrock mode, and account
entitlement; on any miss it silently falls back to the static `models`
array
> - Claude Opus 5 (`claude-opus-5`) is generally available — Anthropic
lists it as the recommended model for complex agentic coding and
enterprise work — but it was absent from that static fallback, so it
appeared only when live discovery happened to succeed
> - This pull request adds `claude-opus-5` to the `claude_local` static
model list so it is selectable regardless of the live-discovery path
> - The benefit is a consistent, reliable dropdown that surfaces the
current GA Opus flagship without depending on a flaky live fetch

## Linked Issues or Issue Description

No public GitHub issue. The bug is described inline following the
bug-report template:

**What happened**
The `claude_local` agent-config model dropdown omitted Claude Opus 5.
`claude-opus-5` was missing from the adapter's static fallback `models`
array (`packages/adapters/claude-local/src/index.ts`), so it only
surfaced when the live Anthropic `/v1/models` discovery happened to
succeed.

**Expected behavior**
Claude Opus 5 is a shipped, generally-available flagship (Anthropic's
recommended model for agentic coding) and should always be selectable in
the dropdown, independent of whether live discovery succeeds.

**Steps to reproduce**
1. Run the server without a working live Anthropic `/v1/models` path (no
`ANTHROPIC_API_KEY`, Bedrock mode, a discovery timeout, or a cache
miss).
2. Open agent config for a `claude_local` agent and inspect the model
dropdown.
3. Observe that `claude-opus-5` is absent because the static fallback
list omitted it.

**Deployment mode**
Self-hosted / local adapter (`claude_local`); the server process reads
`ANTHROPIC_API_KEY` from its environment.

## What Changed

- Added `{ id: "claude-opus-5", label: "Claude Opus 5" }` to the
`claude_local` static `models` fallback. Placed after the current
5-family entries and above the legacy `claude-opus-4-7`, so
`claude-opus-4-8` stays the default (index 0) option.
- Added an explicit regression assertion in
`server/src/__tests__/adapter-models.test.ts` that `claude-opus-5` is
present in the `claude_local` fallback when live discovery is
unavailable.

## Verification

- `pnpm -C server exec vitest run src/__tests__/adapter-models.test.ts`
— **17/17 pass**, including the new `claude-opus-5` assertion and the
existing `models[0] === "claude-opus-4-8"` default invariant (unaffected
— Opus 5 is inserted lower in the list).
- Change is a single static-data addition plus a test assertion; no
control-flow change.

## Risks

- Low risk. Pure additive change to a fallback list; no control-flow
change. Worst case is an id a given account isn't entitled to, which the
existing "current"/manual-model UI paths already tolerate.
- Note for reviewers: a sibling PR adds `claude-sonnet-5` to the same
static array (near `claude-opus-4-8`). Both are complementary "refresh
the static list to current GA" changes; whichever merges second may need
a one-line merge resolution in
`packages/adapters/claude-local/src/index.ts` and the matching test
assertion block.

## Model Used

Claude (Anthropic), model id `claude-opus-4-8` (Opus 4.8), extended
thinking + tool use, run as the Paperclip CTO agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (branch is the assigned
execution-workspace branch and cannot be renamed this run)
- [x] I have run tests locally and they pass (server adapter-models
suite, 17/17)
- [x] I have added or updated tests where applicable (explicit
`claude-opus-5` fallback assertion)
- [x] I have updated relevant documentation to reflect my changes (n/a —
no docs reference this list)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending CI)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending review)
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-28 16:25:25 -05:00
Dotta 487e33b8b6
fix(codex-local): resolve GPT-5.6 model metadata at source (#9780)
## Thinking Path

> - Paperclip is the open source control plane people use to manage AI
agents for work
> - The `codex_local` adapter runs OpenAI's Codex CLI through direct CLI
and ACP execution lanes
> - The adapter defaulted to the bare `gpt-5.6` alias while the bundled
ACP Codex version lacked GPT-5.6-family metadata
> - Default and legacy-configured runs therefore emitted
fallback-metadata warnings and could use generic context limits
> - This pull request upgrades the bundled Codex ACP dependency, selects
the concrete `gpt-5.6-sol` model, and normalizes the legacy alias in
both execution lanes
> - The benefit is correct model metadata without hiding genuine stderr
or transcript warnings

## Linked Issues or Issue Description

Related public PRs: Refs #9342, Refs #9352, and Refs #9382. This PR is
narrower: it upgrades bundled Codex metadata and normalizes the legacy
bare alias in both execution lanes.

**Bug report**

### What happened

Default `codex_local` runs, and agents still configured with the bare
`gpt-5.6` model, print a model-metadata fallback warning and use generic
context-window limits.

Root cause: the ACP lane bundled a Codex release predating
GPT-5.6-family metadata, while Paperclip's default and advertised model
used the bare `gpt-5.6` alias for which Codex publishes no metadata.

### Expected behavior

A default Codex run resolves to a concrete model slug with published
metadata and does not emit a fallback-metadata warning.

### Deployment mode

Self-hosted/local `codex_local` adapter.

## What Changed

- Upgraded `@agentclientprotocol/codex-acp` from `^1.1.0` to `^1.1.4`
- Changed `DEFAULT_CODEX_LOCAL_MODEL` from `gpt-5.6` to `gpt-5.6-sol`
- Removed the bare alias from advertised models and listed concrete
GPT-5.6 Fast-mode variants
- Added `normalizeCodexModel()` and applied it in both CLI and ACP
execution lanes
- Updated adapter docs, Storybook fixtures, and regression tests
- Preserved warning visibility; no stderr, transcript, or log filtering
changed

## Verification

- `pnpm --filter @paperclipai/adapter-codex-local typecheck`
- `pnpm check:token-gates`
- `cd packages/adapters/codex-local && pnpm exec vitest run` — 205 tests
passed
- `cd server && pnpm exec vitest run
src/__tests__/adapter-models.test.ts` — 17 tests passed
- Confirmed the PR diff excludes `pnpm-lock.yaml` and
`.github/workflows/**` as required by repository policy
- Confirmed `.github/workflows/pr.yml` regenerates and uploads the PR
lockfile artifact before downstream `pnpm install --frozen-lockfile`
steps

## Risks

Low risk. The behavior change is scoped to `codex_local` model
selection. Existing concrete model IDs pass through unchanged; only the
legacy bare `gpt-5.6` alias is rewritten. Dependency resolution may
select a newer compatible `codex-acp` release within the declared range,
so CI remains the final compatibility gate.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Original implementation: Anthropic Claude Opus 4.8 (`claude-opus-4-8`,
1M context, tool use and code execution)
- Conflict resolution and PR preparation: OpenAI GPT-5.5 (`gpt-5.5`,
Codex CLI coding agent, high-reasoning tool use and code execution;
host-managed context window)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
— branch name is fixed by the assigned execution workspace and cannot be
renamed in-place
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-28 16:15:08 -05:00
dependabot[bot] f9034ab3ca
build(deps-dev): bump @types/node from 22.19.21 to 22.20.1 (#10304)
Bumps
[@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node)
from 22.19.21 to 22.20.1.
<details>
<summary>Commits</summary>
<ul>
<li>See full diff in <a
href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-28 10:48:45 -07:00
dependabot[bot] 4881f548be
build(deps): bump @cursor/sdk from 1.0.19 to 1.0.24 (#10298)
Bumps [@cursor/sdk](https://github.com/cursor/cursor) from 1.0.19 to
1.0.24.
<details>
<summary>Commits</summary>
<ul>
<li>See full diff in <a
href="https://github.com/cursor/cursor/commits">compare view</a></li>
</ul>
</details>
<details>
<summary>Maintainer changes</summary>
<p>This version was pushed to npm by <a
href="https://www.npmjs.com/~luist18">luist18</a>, a new releaser for
<code>@​cursor/sdk</code> since your current version.</p>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=@cursor/sdk&package-manager=npm_and_yarn&previous-version=1.0.19&new-version=1.0.24)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-28 10:03:36 -07:00
Nicky Leach 341993ebae
feat(sandbox-runtime): route all inbound staging through client.syncIn (Codex home -> native uploadFiles; delete usesCustomProvision gate) (#10354) 2026-07-28 07:22:18 -07:00
Dotta 1426494ab8
fix(agents): disable cheap model profiles by default (#10019)
## Thinking Path

> - Paperclip is the control plane people use to create and govern
AI-agent companies
> - Agent creation persists runtime configuration that controls which
model profiles future runs may select
> - Adapters can expose a `cheap` profile, and existing creation paths
implicitly left that profile available when operators made no choice
> - That made a newly created agent eligible for a lower-cost model
without an explicit operator opt-in
> - The UI also dropped an explicit opt-in when the operator selected
the adapter's default cheap model rather than a custom model ID
> - Codex additionally hardcoded `gpt-5.3-codex-spark` into its cheap
profile and static fallback model list, making Paperclip choose an
auth-dependent model rather than requiring an operator choice
> - This pull request makes new-agent creation disable an available
cheap profile by default while preserving explicit opt-in from the UI or
API
> - The Codex cheap profile now remains available for explicit
configuration but supplies no model default, so an unconfigured cheap
request stays on the primary model
> - The benefit is predictable model quality for new agents and an
intentional, auditable choice before lower-cost routing is enabled

## Linked Issues or Issue Description

**Problem**

New agents created with an adapter that exposes a `cheap` model profile
can inherit that profile without the operator explicitly enabling it. In
the UI, enabling the adapter-default cheap model is also omitted because
runtime configuration is only written when a custom model ID is present.

**Expected behavior**

- New agents default an available `cheap` model profile to `{ enabled:
false }` when the caller does not specify it.
- Explicit API configuration remains authoritative.
- UI opt-in persists even when the adapter default model is used.
- Codex does not advertise or automatically select
`gpt-5.3-codex-spark`; operators must explicitly configure any
lower-cost Codex model.

**Related public work**

- Refs #4881, which introduced cheap model profiles for local adapters.
- Supersedes the default-selection portions of #8032 and #10004 by
removing the Codex model default instead of replacing it with another
hardcoded model.

## What Changed

- Detect whether the selected adapter exposes a `cheap` model profile
during agent creation and hiring.
- Persist `runtimeConfig.modelProfiles.cheap.enabled = false` only when
the caller did not explicitly configure the profile.
- Preserve UI cheap-profile opt-in when using the adapter's default
model by writing an empty adapter config.
- Remove `gpt-5.3-codex-spark` from the Codex static model list.
- Keep the Codex `cheap` profile explicitly configurable while giving it
an empty adapter config, so Paperclip never chooses a cheap Codex model
automatically.
- Verify that a Codex cheap request without an explicit model leaves the
primary model unchanged.
- Extend server route and UI runtime-config tests for default-disable
and explicit-opt-in behavior.

## Verification

- `env -u PAPERCLIP_IN_WORKTREE -u PAPERCLIP_WORKTREE_NAME -u
PAPERCLIP_CONFIG -u PAPERCLIP_HOME -u PAPERCLIP_INSTANCE_ID -u
PAPERCLIP_CONTEXT pnpm exec vitest run
packages/adapters/codex-local/src/index.test.ts
packages/adapters/codex-local/src/server/codex-args.test.ts
server/src/__tests__/adapter-models.test.ts
server/src/__tests__/adapter-registry.test.ts
server/src/__tests__/heartbeat-model-profile.test.ts
server/src/__tests__/agent-permissions-routes.test.ts
ui/src/lib/new-agent-runtime-config.test.ts`
- Result: 7 test files passed, 105 tests passed.
- GitHub `Typecheck + Release Registry` check passed on the final head.
- `git diff --check public-gh/master...HEAD`

## Risks

- Low behavioral risk: only newly created or hired agents are
normalized; existing agents are unchanged.
- Explicit `cheap` profile settings remain untouched, including explicit
opt-in.
- Codex users who explicitly opt into the cheap lane must choose a
model; requests without a configured override intentionally continue on
the primary model.
- Adapter profile discovery is now awaited during creation, adding a
small amount of adapter metadata lookup work.
- The source branch name is automation-provided and retained as required
by the task, so it does not satisfy the preferred public branch naming
convention.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI `gpt-5.4` via Codex CLI, with reasoning, repository editing,
terminal execution, and GitHub/Paperclip tool access. The runtime did
not expose a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-27 19:11:51 -05:00