A suffix segment adjacent to a serialized quoted header argument belongs to
the same shell word, so the serialized branch now takes the shell
continuation after its closer. A truncated serialized argument has no
closer on its line; its value then stops before a bare quote, one preceded
by an even run of backslashes, which can only be the enclosing serializer's
delimiter, so that string stays parseable. The closer of an escaped-quoted
value must itself be unescaped, so backtracking cannot read an escaped
backslash as the closer.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
An escaped-quoted argument carries a backslash run before its quote that
doubles and grows by one with every serialization layer, so a rule that
requires exactly one backslash misses a command serialized twice. Both
escaped-quote branches now capture the odd backslash run of the opener and
close on the same run, with a tempered body that keeps deeper embedded
quotes and a dangling trailing backslash inside the value. A scheme word
may precede the escaped opener as well as follow it. Every capture group
is named and the replace callback reads the groups object.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
A value written with escaped quotes after the colon, the form an outer
shell uses to pass quote syntax to `sh -c`, had no owner: the unquoted
branch declines an escaped-quote opener so the server's own authorization
rule keeps its shape. A dedicated branch now redacts that value and keeps
the escaped quotes, so both rules agree on the same text. The unquoted
branch also keeps a value's own delimiters when the value is quoted after
the colon, which makes the rule idempotent across the log-writer and UI
passes. The replace callback selects prefix, opener, and closer from the
defined capture groups.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
The first segment of an unquoted header value may open on an escape pair,
so a value whose first byte is an escaped space is consumed, while an
escaped-quote opener still falls to the caller's own rules. That first
segment is bounded by whitespace only: a raw HTTP diagnostic carries an
opaque credential the same way, so a shell metacharacter inside it is a
credential byte. Only a continuation segment after a closing quote stops
at a metacharacter, which keeps a following separator or command intact.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
A header value is the rest of its shell word, which can concatenate
unquoted, double-quoted, single-quoted, ANSI-C-quoted, and backslash-escaped
segments. The rule now consumes every segment of that word before writing
one placeholder, stops at whitespace and shell metacharacters so the next
argument survives, and still redacts a run-log line truncated inside a
quoted value. The recognized scheme list gains the registered `Concealed`
scheme, and a quoted Digest parameter may carry HTTP quoted-pairs.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
A serialized command writes a double-quoted header argument with escaped
quotes. The escaped opener fell through to the unquoted branch, which
stops at the first backslash, so a multi-part credential such as a Digest
value kept its later fields. A fourth branch mirrors the double-quoted one
over `\"` delimiters and consumes the doubled escape sequences an embedded
quote or backslash becomes.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
A backslash-newline continuation inside a double-quoted header argument
ended the quoted match, so the unquoted fallback redacted only the part of
the credential before the continuation. The double-quoted branch now treats
the continuation as part of the value, with LF and CRLF line endings.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
A double-quoted header value stopped at the first backslash, so a
credential with an embedded escaped quote such as `"X-API-Key: abc\"def"`
kept its tail in the recorded text. The double-quoted branch now consumes
escape pairs and requires an unescaped opening quote, which keeps it off a
serialized diagnostic where `\"` is the JSON escape. The single-quoted
branch takes a backslash literally. Only the unquoted branch still stops at
a backslash.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
The header rule stopped at the first whitespace or quote, so a multi-part
credential such as a Digest or AWS SigV4 authorization value lost only its
first token. It also matched any bare word carrying a credential hint, so
prose and paths like `auth: failed` or `/v1/tokens:list` were redacted.
The value is now bounded by its context: to the closing quote inside a
quoted shell argument, and to the end of a comma-separated `key=value` list
or a single token when unquoted. A header name must be hyphenated or
underscored, or be the bare `authorization` or `apikey`; the
`www-authenticate` and `proxy-authenticate` challenge headers are excluded.
The recognized scheme list follows the IANA registry plus
`AWS4-HMAC-SHA256` and `Token`.
Claude-Session: https://claude.ai/code/session_01RYigf3eMFJjey9iKRApPGE
The command redaction covered `Authorization: Bearer <value>`, shell
`NAME=value` assignments, and common token shapes. It did not cover a
credential passed in any other header. A `curl -H "X-API-Key: <token>"`
command therefore kept the token in clear in a run log.
A new rule redacts the value of any header whose name contains an api-key,
token, secret, or auth hint. The rule keeps an optional auth scheme in the
output, so `Authorization: Bearer <value>` produces the same text as
before. `Authorization: Basic <value>` is now redacted too. The value ends
at the first quote, backslash, or whitespace, so the rule stops at the end
of one header argument.
Claude-Session: https://claude.ai/code/session_01U9PF3d9SASC9tomDRjyeVt
## Thinking Path
> - Paperclip is the control plane for agents that perform work.
> - Paperclip Runner connects durable provider sessions to individual
task runs through PRP.
> - Provider continuity and per-run authority are different lifetimes.
> - The existing implementation mixed those lifetimes and lost event
metadata between provider frames, runnerd, persistence, API
sanitization, and the task thread.
> - That caused failed continuation, missing progress and Plans,
duplicate replies, hidden failures, and unsafe recovery.
> - This repair gives every heartbeat fresh authority, preserves
qualified provider-session continuity, and restores one lossless
presentation path without changing direct adapters.
## Linked Issues or Issue Description
**What happened?**
A second native heartbeat could reuse tickets, leases, command receipts,
sequence state, and run identity from the first heartbeat. Provider
phase and item identity could be lost before the UI read them. Redaction
could corrupt protocol discriminators while still missing malformed
credential tails. The task thread could fold progress into the final
response, hide failures, or show more than one final answer. Native
Codex also exposed approval modes that do not yet have a durable
approval bridge.
**Expected behavior**
Each heartbeat uses a new PRP authority epoch. Codex and OpenCode
preserve exact qualified provider sessions; ACPX emits an explicit
continuity event when its qualified process-replacement policy is used.
Every accepted provider event is presented, classified as internal, or
surfaced as unsupported. The task page shows chronological progress,
reasoning summaries, activity, Plans, interactions, terminal failures,
and exactly one final reply. Direct adapters retain their existing path.
**Steps to reproduce**
1. Enable the unified experimental Paperclip Runner setting.
2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex
agent.
3. Run response, Plan, structured-question/resume, restart,
cancellation, and failure scenarios.
4. Reload the task while active, waiting, failed, and settled.
5. On the old implementation, observe stale run authority, missing
classifications, incomplete output, or duplicated/folded replies.
**Paperclip version or commit**
The repair is based directly on `master` at
`87d05e194b643810d16d20612115acd01d735d43`.
**Deployment mode**
Local development with the embedded database.
Related work: Refs #12616, #12646, #12666, #12685, and #12700.
## What Changed
- Rotates PRP control-plane, outbox, ticket, lease, command, receipt,
and sequence authority for each heartbeat while carrying forward only a
validated provider-session identity.
- Reads `control-plane-state.json`, validates both durable schemas and
lifecycle values, resumes coherent current runs, archives qualified
settled authority, and quarantines malformed or mismatched scoped state
without moving ambiguous live legacy state.
- Preserves Codex provider phase and stable item identities so
commentary remains progress and only `final_answer` becomes final.
- Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning
lifecycle mapping.
- Makes ACPX normalization lossless for visible reasoning, tool
lifecycle metadata, stable bounded identities, Plan revisions,
structured requests, failures, and qualified process replacement. Only
the compatible terminal assistant message is promoted as final.
- Applies schema-aware redaction before generic JWT-shaped detection and
scans every diagnostic string leaf. Malformed raw/escaped quoted
credential tails are redacted in both server and durable Rust state.
- Restores snapshot-style chronological task presentation, expandable
tool activity, inline Plan cards, visible waiting/resume/cancel/failure
states, and exactly one final answer.
- Makes `never` the only qualified native Codex permission mode and
rejects unsupported persisted native modes with remediation. OpenCode
and ACPX policies remain intact.
- Keeps the unified experimental Runner setting as the only enablement
flag. Onboarding and direct Codex, Claude, and OpenCode stay on their
legacy execution/finalization paths.
- Adds cross-language goldens, authority/recovery/fault coverage, exact
response/count assertions, and native plus legacy acceptance scenarios.
## Verification
- Pull-request GitHub Actions run Rust formatting/tests, TypeScript
checks, server/UI tests, builds, protocol drift checks, browser E2E, and
security scans.
- A separate workflow-only validation ref is pinned directly on this PR
head and runs the 35-cell paid local matrix: three core scenarios plus
structured-question resume and restart/resume for native Codex, native
OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode.
Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315
- Acceptance requires exact single visible replies, monotonic sequences,
matching envelope discriminators, one semantic terminal, one run
terminal, no unresolved interaction, no duplicate mutation, no secret
leakage, provider continuity, and zero native rows for direct adapters.
- Per maintainer direction, tests are running in GitHub Actions rather
than on the slower local host. Only formatters and static diff checks
were run locally.
## Risks
- Recovery from old or partial filesystem state is sensitive. The repair
fails closed, preserves active or unverifiable authority, and
quarantines only state whose scoped ownership is safe to move.
- Provider event formats can change. Closed validators and boundary
goldens turn new or malformed events into visible diagnostics instead of
silent drops.
- Shared task presentation could affect direct adapters. Runtime-fact
gating plus the direct-adapter matrix protect the existing path.
- Managed and remote providers are not qualified here. Shared code
continues to compile and fail safely, but live qualification is
deferred.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex based on GPT-5. The exact deployed snapshot and
context-window size are not exposed to this task. It used agentic
reasoning, repository inspection, code editing, Git, parallel subagents,
and GitHub Actions.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] The paid local-provider matrix is green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The Claude local adapter supports subscription login through a
sandbox
> - The new-agent page must show login before the user creates an agent
> - Test results must not expose raw sandbox diagnostics or secret
values
> - This pull request adds the login UI to both Test lanes and closes
the diagnostic boundary
> - The branch also adds durable cleanup recovery for failed sandbox
teardown
> - Reusable sandboxes must retain both their recorded teardown
configuration and a valid lifecycle path until destruction succeeds
> - The benefit is a usable login flow with fixed public checks,
redacted server logs, and recoverable sandbox cleanup
## Linked Issues or Issue Description
Related public work:
[#9488](https://github.com/paperclipai/paperclip/pull/9488) adds
first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and
remote runs. Related public issue:
[#2681](https://github.com/paperclipai/paperclip/issues/2681) requests
Claude Code subscription support. This pull request adds the login
transport and new-agent UI flow that those changes do not provide.
**Subsystem affected:** Claude local adapter, server login probes,
sandbox provider setup, cleanup recovery, and the new-agent UI.
**Problem or motivation:** The Test lanes did not show the sandbox login
panel in all supported cases. Test results also exposed raw probe
diagnostics, and JSON escapes could end secret redaction early.
**Proposed solution:** Surface the login capability through the bundled
provider manifest. Prepare the same probe runtime in the ACP lane. Send
diagnostics only to redacted server logs. Keep Test checks on fixed
public messages. Normalize login URL hints to allowlisted HTTPS Claude
and Anthropic hosts. Consume JSON escapes during redaction. Preserve
failed sandbox cleanup state across retries and restarts, and prevent
deletion from severing the lifecycle context of a live reusable sandbox.
**Alternatives considered:** Keep raw diagnostics in Test checks or
trust login URL text from the sandbox. Both choices increase information
exposure. Keep separate probe behavior in the ACP lane. That choice
would leave the two Test lanes inconsistent.
## What Changed
- Surface the sandbox login panel on both Test lanes.
- Reconcile the bundled Daytona plugin manifest so
`supportsSetupTokenLogin` reaches the UI capability gate.
- Prepare the ACP Test lane with the same probe runtime as the CLI Test
lane.
- Add the `claude_acp_login_probe_unavailable` warning when the ACP
probe cannot run.
- Send raw sandbox diagnostics only to redacted server logs.
- Keep Test checks on fixed public messages in the ACP, managed-config,
and CLI paths.
- Normalize login URL hints to allowlisted HTTPS Claude and Anthropic
hosts.
- Redact JSON and escaped-JSON secret values, including escaped quotes
and backslashes.
- Preserve orphan cleanup records across provider failures, restarts,
and unavailable plugins.
- Atomically block environment deletion while a live reusable sandbox
lease still depends on it.
- Verify pending cleanup destroys plugin sandboxes with the provider
configuration recorded on the lease, even after the current environment
configuration changes.
## Verification
- Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`.
- Focused environment route/service/runtime coverage passes: 196 tests
across 3 files.
- `pnpm -r typecheck` passes.
- `pnpm build` passes.
- The full Vitest run completed with 4,754 passing and 28 failing tests.
All 23 source-test failures reproduce unchanged on parent head
`58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate
executions from stale `server/dist` output. The failures are unrelated
macOS path/listener and scheduler-fixture failures, so there is no new
bad commit for bisect to localize.
- All required CI checks pass for the current head, including build,
typecheck/release registry, all server and workspace shards, serialized
server suites, canary, and e2e.
- A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`
reports 5/5, “safe to merge,” with no blocking failure remaining.
## Risks
- A probe or redaction change could hide useful server diagnostics.
- An allowlist change could reject a valid Claude login URL.
- Cleanup recovery changes could affect provider teardown ordering.
- An environment with a live reusable sandbox can no longer be deleted
until the owning issue or execution workspace completes teardown.
- The implementation keeps public Test messages fixed and sends detail
to redacted server logs.
## Model Used
OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code
execution enabled; extended reasoning enabled. The implementation author
used AI-assisted development.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and documented the result
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation or confirmed no separate
documentation change is needed
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>