paperclip/packages/paperclip-runner
Dotta 42b8f7ab2f
feat(runner): authorize semantic tool dispatch (#12126)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package defines a provider-neutral protocol and semantic
action catalog.
> - Catalog membership alone must not grant access to an action.
> - Each run needs current company, actor, task, claim, mode, and
application-binding authority.
> - Mutating actions also need safe retry behavior and durable receipts.
> - This pull request adds a package-local authority and dispatch layer.
> - The benefit is a small and testable trust boundary before server
integration lands.

## Linked Issues or Issue Description

Refs #11962

This pull request replaces one bounded part of the archived large runner
change.

## What Changed

- Add run-scoped tool projection and optional tool discovery.
- Require an explicit application binding before an action is visible.
- Intersect actor claims with claims delegated to the run.
- Recheck company, actor, task, mode, state, role, claim, and policy
authority before each call.
- Validate action input and output with the canonical catalog schemas.
- Redact protected values and keep raw tool content out of semantic
receipts.
- Require atomic idempotency claims for mutating actions.
- Replay exact completed retries and reject changed or concurrent
retries.
- Recover a durable completed receipt if the primary receipt commit
fails, without re-executing the mutation.
- Add bounded authorization records and PRP semantic input and result
receipts.
- Document that this change adds no server binding or production tool
installation.

## Verification

- `pnpm --filter @paperclipai/paperclip-runner check:all`
- `pnpm -r typecheck`
- `pnpm check:token-gates`
- `pnpm build`
- 60 package TypeScript tests pass.
- 56 Rust unit and integration tests pass.
- Protocol, replay, and cross-language conformance checks pass.
- `pnpm test:run` completed with 4,684 passing and 19 skipped tests. It
reproduced 32 local baseline failures across 9 unchanged server files;
all corresponding hosted test shards pass.
- Every applicable GitHub Actions gate passes. The Storybook job skipped
because this PR has no UI changes.
- Socket and Snyk pass with no findings. Superagent completed neutral
with zero annotations because its external sandbox did not start within
120 seconds.
- Greptile is 5/5 with no unresolved actionable comments.
- The diff changes 11 files.

## Risks

The main risk is an authorization or idempotency error at the tool
boundary. The dispatcher fails closed for malformed authority,
unavailable receipt storage, stale authority, unauthorized actions,
protected input, invalid binding output, and unrecoverable receipt
completion. The receipt store must recover a completed mutation outcome
idempotently if its primary commit fails; otherwise the claim remains
reserved for operator recovery rather than allowing automated
re-execution. Unbound actions are absent. No server or provider installs
these tools in this change. Existing adapters and application behavior
do not change.

## Model Used

OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code
execution, and automated tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the affected tests locally and they pass; full-suite
baseline exceptions are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-24 17:28:54 -05:00
..
generated Add canonical semantic action catalog to Paperclip Runner (#12121) 2026-08-24 16:26:21 -05:00
protocol Add local fake runner supervision (#12095) 2026-08-24 12:16:48 -05:00
runner Add Codex provider bridge to Paperclip Runner (#12111) 2026-08-24 15:19:14 -05:00
scripts Add canonical semantic action catalog to Paperclip Runner (#12121) 2026-08-24 16:26:21 -05:00
src feat(runner): authorize semantic tool dispatch (#12126) 2026-08-24 17:28:54 -05:00
test feat(runner): add PRP v1 schemas and fixtures (#12087) 2026-08-24 09:59:03 -05:00
.gitignore Add local fake runner supervision (#12095) 2026-08-24 12:16:48 -05:00
README.md feat(runner): authorize semantic tool dispatch (#12126) 2026-08-24 17:28:54 -05:00
SEMANTIC_ACTIONS.md feat(runner): authorize semantic tool dispatch (#12126) 2026-08-24 17:28:54 -05:00
package.json Add canonical semantic action catalog to Paperclip Runner (#12121) 2026-08-24 16:26:21 -05:00
tsconfig.json Add TypeScript PRP replay contracts (#12091) 2026-08-24 10:43:53 -05:00
vitest.config.ts Add TypeScript PRP replay contracts (#12091) 2026-08-24 10:43:53 -05:00

README.md

Paperclip Runner

This private workspace package contains the staged Paperclip Runner work.

The package currently exposes only the language-neutral PRP v1 TypeScript contract, provider-neutral structured questions and responses, deterministic fixture validation/replay, structured-result normalization, and the session reducer oracle. It also contains a package-local Rust runner, scripted fake harness, bounded process supervisor, cross-language replay oracle, and durable PRP transport. The transport authenticates and encrypts loopback WebSocket sessions, persists an ACK-driven outbox and command journal, and reconnects with a short-lived lease. The Rust runner now includes a Codex-only app-server provider bridge with durable thread resume, cancellation, structured questions, and provider-neutral event normalization. No server code starts or invokes it. The package also publishes the canonical semantic action declarations and their input and output schemas. Its package-local dispatcher projects only bound, run-authorized actions and emits redacted semantic receipts. It does not add application bindings, a server adapter, or production Paperclip behavior.

The first and only installed provider is Codex. Dynamic semantic tools remain undiscoverable because no production application binding or server authority has landed. Catalog membership alone does not grant authority. See SEMANTIC_ACTIONS.md for the catalog boundary.

The root export is intentionally narrow. The ./testing entry point and package release boundary will arrive with the later package-boundary change.

Run the complete contract gate with:

pnpm --filter @paperclipai/paperclip-runner check:protocol

Run the Rust runner gate with:

pnpm --filter @paperclipai/paperclip-runner check:runner

This command checks Rust formatting, builds and tests the minimal workspace, verifies bounded process cleanup, exercises the fake local runner, and compares the Rust conformance and replay summaries with the shared fixtures.

Durability and failure semantics are documented in runner/DURABLE_TRANSPORT.md. The fault suite drops a connection before its event ACK, reconnects with the bound lease, replays the same event, and proves the duplicated command effect ran once. Codex launch, resume, cancellation, and normalization behavior is documented in runner/CODEX_PROVIDER.md.

Use generate:protocol-manifest after a schema or fixture change, generate:protocol-types after a schema change, and generate:replay-goldens after an intentional reducer change. Commit generated outputs with their sources; do not edit them by hand.

Use generate:semantic-action-catalog after changing a semantic action declaration. Its checked-in JSON inventory must land with the source change.

The gate compiles every schema with AJV 2020-12, validates accepted fixtures, rejects unsupported required versions, checks generated TypeScript schema drift, runs the TypeScript contract tests, and compares reducer snapshots and parity summaries byte-for-byte with their checked-in golden files.