paperclip/packages/paperclip-runner/spec/paperclip-agent-operation-g...

82 KiB

Paperclip agent operation groups

Status: canonical explanatory contract for the Paperclip runner V1 surface.

This document keeps three independent meanings of group separate. PRP families describe wire evidence and controller commands; capability placement decides who owns an operation; behavioral eval groups organize the 106 scenario corpus. None of the three axes can be used as a substitute for another.

The generated totals are 105 PRP events in 31 event families, 18 controller commands in 7 command families, 10 control-plane operations, 41 reconciled semantic operations (14 always, 27 optional), and 106 scenarios in 16 behavior groups.

Axis 1: PRP v1 event and command families

PRP records ordered, replayable execution evidence. It is not the model's Paperclip tool API. Events are ordered per sourceInstanceId; duplicate sourceEventId values are idempotent; source-sequence gaps remain evidence; replay is side-effect free; unknown required versions fail closed.

Event families

Family Purpose Events Count
runner Runner connection, drain, and diagnostic lifecycle. runner.connected
runner.reconnected
runner.reconciled
runner.disconnected
runner.draining
runner.backpressure
runner.suspending
runner.suspended
runner.stopped
runner.diagnostic
10
runtime Runner phase transitions. runtime.phase.changed 1
sandbox Sandbox resource measurements. sandbox.metric 1
workspace Workspace readiness. workspace.ready
workspace.change.updated
workspace.diff.recorded
workspace.file.referenced
4
harness Provider harness startup, readiness, exit, and diagnostics. harness.starting
harness.ready
harness.exited
harness.diagnostic
4
plan Complete provider-authored within-turn checklist snapshots, separate from durable Paperclip Plan documents. plan.updated 1
tool Provider-neutral process, MCP, dynamic, and built-in execution activity. tool.execution.started
tool.execution.progressed
tool.execution.completed
3
research Provider-reported search, page-open, and in-page research activity. research.started
research.progressed
research.completed
3
delegation Child-agent delegation lifecycle and aggregate status. delegation.started
delegation.updated
delegation.completed
3
model Requested/effective model routing and verification state. model.route.changed
model.verification.updated
2
context Context-window compaction markers without hidden summaries. context.compacted 1
artifact Authorized artifact viewing and structured generated outputs. artifact.viewed
artifact.generated
2
review Provider review-mode state, separate from Paperclip authority. review.mode.changed 1
hook Bounded provider hook lifecycle and blocking outcomes. hook.started
hook.completed
2
memory Authorized or unavailable memory citation references. memory.citation.referenced 1
safety Provider safety review state attached to governed work. safety.review.started
safety.review.completed
2
terminal Content-free terminal input activity metadata. terminal.input.sent 1
wait Intentional provider waits, distinct from warm idle and human input. wait.started
wait.completed
2
provider Redacted provider notices and actionable warnings. provider.notice.recorded 1
session Provider-neutral session open, resume, reconciliation, close, and failure. session.starting
session.started
session.resuming
session.resumed
session.reconciled
session.updated
session.closed
session.failed
8
turn Model turn submission through terminal turn disposition. turn.submitted
turn.accepted
turn.started
turn.completed
turn.failed
turn.interrupted
turn.cancelled
7
item Provider-neutral model/tool item lifecycle. item.started
item.delta
item.completed
item.failed
4
usage Provider/model-attributed usage and accounting boundaries. usage.reported 1
semantic_tool Canonical authorized Paperclip tool input and result evidence. semantic_tool.input
semantic_tool.result
semantic_tool.reconciled
3
mcp_app MCP App discovery, initialization, tool, action, host-context, and teardown evidence. mcp_app.discovered
mcp_app.resource.resolved
mcp_app.initializing
mcp_app.ready
mcp_app.tool_input
mcp_app.tool_result
mcp_app.action.requested
mcp_app.action.resolved
mcp_app.host_context.changed
mcp_app.failed
mcp_app.teardown
11
runtime_request Runtime permission/input request lifecycle. runtime_request.created
runtime_request.resolved
runtime_request.expired
runtime_request.cancelled
4
interaction Issue-thread interaction proposal, materialization, response, delivery, and rejection. interaction.request.proposed
interaction.request.materialized
interaction.request.rejected
interaction.response.progressed
interaction.response.resolved
interaction.response.delivered
6
run Structured result negotiation and terminal run outcome. run.attached
run.detached
run.result.proposed
run.result.accepted
run.result.rejected
run.terminal
6
attention Continuation and attention routing lifecycle. attention.request.proposed
attention.request.routed
attention.request.resolved
attention.request.expired
attention.request.superseded
5
work Recorded work-assessment evidence. work.assessment.recorded 1
issue Issue-status decision proposal outcome and application evidence. issue.status.decision.recorded
issue.status.decision.applied
issue.status.decision.rejected
issue.status.decision.superseded
4

Controller-command families

Family Purpose Commands Count
run Prepare or cancel a run. run.prepare
run.attach
run.cancel
3
session Open, snapshot, or close a normalized provider session. session.open
session.snapshot
session.close
session.budget.increase
session.destroy
5
turn Start, steer, interrupt, or stop a model turn. turn.start
turn.steer
turn.interrupt
turn.stop
4
request Resolve a pending runtime request. request.resolve 1
interaction Acknowledge delivery of an interaction response. interaction.receipt 1
semantic_tool Returns an authorized, correlated Paperclip tool result to the provider through runnerd. semantic_tool.result 1
runner Drain or shut down the runner process. runner.drain
runner.suspend
runner.shutdown
3

Axis 2: capability placement

Placement has exactly three outcomes:

  • control_plane_owned: Paperclip or the runner performs the operation; it is absent from the model tool catalog.
  • always_agent_tool: every eligible active-task run receives the operation after task-mode and actor checks.
  • optional_agent_tool: the operation is exposed only when every declared claim, task-mode, role, and policy condition passes.

Control-plane-owned operations

Operation Why it is not a model tool
checkout_task Atomic checkout and execution-lock ownership.
release_task Run cleanup and checkout release.
select_work Scoped wake/inbox work selection.
route_wake Attention, blocker, interaction, approval, and continuation routing.
enforce_budget Budget hard stop, pause, and run-stop reason.
append_audit_record Immutable mutation evidence.
persist_run Durable run/event/checkpoint persistence.
replay_run Side-effect-free replay and reconstruction.
schedule_blocker_wake Dependency-resolution wake scheduling.
reconcile_run Work assessment, terminal status arbitration, and recovery reconciliation.

Always-agent operations (14)

answer_status_question, block_task, finish_task, get_task_context, get_task_history, inspect_operation_result, list_document_revisions, list_documents, read_document, register_deliverable, report_progress, request_human_input, request_review, write_document.

Optional operations (27) and grant groups (11)

Grant groups are documentation/exposure bundles, not additional authority. The operation descriptor's exact requiredClaims remains decisive.

Grant group Operations Required claims represented Purpose
discovery search_tasks
list_agents
get_agent
list_projects
list_goals
discovery:agents:read
discovery:goals:read
discovery:projects:read
discovery:tasks:read
Company-visible task, agent, project, and goal discovery.
delegation_dependencies create_task
set_dependencies
delegation:tasks:create
dependencies:write
Create delegated work and maintain dependency edges.
governance list_approvals
get_approval
get_approval_context
request_approval
decide_approval
comment_on_approval
governance:approvals:comment
governance:approvals:decide
governance:approvals:read
governance:approvals:request
Read, request, comment on, and decide approvals under governed-action checks.
cases list_cases
upsert_case
cases:read
cases:write
Read and update case summaries without reusing issue-document authority.
workspace_runtime get_workspace_runtime
control_workspace_service
workspace:control
workspace:read
Inspect and control the active issue workspace runtime.
wake_scheduling schedule_wake control_plane:wakes Schedule a bounded continuation wake when the current task owns the future check.
routines list_routines
manage_routine
routines:read
routines:write
Inspect or manage company routines.
company_skills list_company_skills
sync_company_skills
company_skills:read
company_skills:write
Inspect company skills and synchronize the current agent's skills.
secrets list_secret_metadata
read_secret_value
secrets:metadata:read
secrets:values:read
Inspect secret metadata or use a brokered secret value without exposing plaintext evidence.
portability_admin export_company
administer_company
company:admin
portability:export
Export portable company state; broad administration remains deferred until split into governed operations.
test_escape_hatch generic_api_request test:generic_api_request Controlled skill-test transport only; never product coverage.

Reconciled semantic-operation ledger

live_codex means a live provider dispatcher exists. Production binding is tracked separately: bound rows are advertised by Paperclip's run-scoped authority over the shared PRP route, while audit_pending rows remain unavailable to production agents. generic_api_request is test-only and cannot satisfy product coverage.

Operation Placement Claims Modes / roles Side effect Idempotency Redacts Mock Catalogs / current runner Production / PRP evidence
administer_company optional_agent_tool company:admin standard
skill_test
roles: board
ceo
admin
admin required no mock_extension:company.admin scenario
scenario_mock
unbound
company admin/portability item event plus audit record
catalog PRP status: audit_pending
answer_status_question always_agent_tool none standard
ask
planning
skill_test
task_write required no semantic_command:report_progress scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
block_task always_agent_tool none standard
skill_test
task_write required no semantic_command:block_task scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
comment_on_approval optional_agent_tool governance:approvals:comment standard
ask
planning
skill_test
governance required no semantic_command:comment_on_approval scenario + live
live_codex
unbound
approval lifecycle plus governed-wait continuation and audit events
catalog PRP status: audit_pending
control_workspace_service optional_agent_tool workspace:control standard
skill_test
workspace_control required no semantic_command:control_workspace_service scenario + live
live_codex
unbound
workspace service lifecycle event
catalog PRP status: audit_pending
create_task optional_agent_tool delegation:tasks:create standard
skill_test
company_write required no semantic_command:create_task scenario + live
live_codex
issues.createChild
semantic-operation item event plus company-entity state diff and audit record
catalog PRP status: bound
decide_approval optional_agent_tool governance:approvals:decide standard
skill_test
roles: board
approver
security
governance required no semantic_command:decide_approval scenario + live
live_codex
unbound
approval lifecycle plus governed-wait continuation and audit events
catalog PRP status: audit_pending
export_company optional_agent_tool portability:export standard
skill_test
admin required no mock_extension:portability.export scenario
scenario_mock
unbound
company admin/portability item event plus audit record
catalog PRP status: audit_pending
finish_task always_agent_tool none standard
skill_test
task_write required no semantic_command:finish_task scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
generic_api_request optional_agent_tool test:generic_api_request skill_test test_escape_hatch required yes mock_extension:test.generic_api scenario + live
test_only
unbound
test-only; excluded from product PRP evidence
catalog PRP status: audit_pending
get_agent optional_agent_tool discovery:agents:read standard
ask
planning
skill_test
read none no inline/no mapping live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
get_approval optional_agent_tool governance:approvals:read standard
ask
planning
skill_test
read none no inline/no mapping live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
get_approval_context optional_agent_tool governance:approvals:read standard
ask
planning
skill_test
read none no inline/no mapping live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
get_task_context always_agent_tool none standard
ask
planning
skill_test
read none no context_read:active_task scenario + live
live_codex
PaperclipRunnerToolAuthority active issue/run + accepted plan revision
bound company/assignment query plus exact accepted document revision projection
catalog PRP status: bound
get_task_history always_agent_tool none standard
ask
planning
skill_test
read none no snapshot_read:active_task_history scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
get_workspace_runtime optional_agent_tool workspace:read standard
ask
planning
skill_test
read none no snapshot_read:active_task_workspace scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
inspect_operation_result always_agent_tool none standard
ask
planning
skill_test
read none no operation_result scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_agents optional_agent_tool discovery:agents:read standard
ask
planning
skill_test
read none no snapshot_read:company_actors scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_approvals optional_agent_tool governance:approvals:read standard
ask
planning
skill_test
read none no snapshot_read:company_approvals scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_cases optional_agent_tool cases:read standard
skill_test
read none no mock_extension:cases.list scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_company_skills optional_agent_tool company_skills:read standard
skill_test
read none no mock_extension:company_skills.list scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_document_revisions always_agent_tool none standard
ask
planning
skill_test
read none no snapshot_read:active_task_document_revisions scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_documents always_agent_tool none standard
ask
planning
skill_test
read none no snapshot_read:active_task_documents scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_goals optional_agent_tool discovery:goals:read standard
skill_test
read none no mock_extension:discovery.goals scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_projects optional_agent_tool discovery:projects:read standard
skill_test
read none no mock_extension:discovery.projects scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_routines optional_agent_tool routines:read standard
skill_test
read none no mock_extension:routines.list scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
list_secret_metadata optional_agent_tool secrets:metadata:read standard
skill_test
read none no mock_extension:secrets.metadata scenario
scenario_mock
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
manage_routine optional_agent_tool routines:write standard
skill_test
admin required no mock_extension:routines.manage scenario
scenario_mock
unbound
company admin/portability item event plus audit record
catalog PRP status: audit_pending
read_document always_agent_tool none standard
ask
planning
skill_test
read none no snapshot_read:active_task_document scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
read_secret_value optional_agent_tool secrets:values:read standard
skill_test
secret_read none yes mock_extension:secrets.value scenario
scenario_mock
unbound
redacted tool-result item event; secret value never reaches the wire
catalog PRP status: audit_pending
register_deliverable always_agent_tool none standard
planning
skill_test
task_write required no semantic_command:register_deliverable scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
report_progress always_agent_tool none standard
ask
planning
skill_test
task_write required no semantic_command:report_progress scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
request_approval optional_agent_tool governance:approvals:request standard
skill_test
governance required no semantic_command:request_approval scenario + live
live_codex
unbound
approval lifecycle plus governed-wait continuation and audit events
catalog PRP status: audit_pending
request_human_input always_agent_tool none standard
planning
ask
skill_test
task_write required no semantic_command:request_human_input scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
request_review always_agent_tool none standard
skill_test
task_write required no semantic_command:request_review scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
schedule_wake optional_agent_tool control_plane:wakes standard
skill_test
task_write required no inline/no mapping live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending
search_tasks optional_agent_tool discovery:tasks:read standard
ask
planning
skill_test
read none no snapshot_read:company_tasks scenario + live
live_codex
unbound
read projection surfaced via a tool-result item event; no control-plane state diff
catalog PRP status: audit_pending
set_dependencies optional_agent_tool dependencies:write standard
skill_test
company_write required no semantic_command:set_dependencies scenario + live
live_codex
issues.update.blockedByIssueIds
semantic-operation item event plus company-entity state diff and audit record
catalog PRP status: bound
sync_company_skills optional_agent_tool company_skills:write standard
skill_test
admin required no mock_extension:company_skills.sync scenario
scenario_mock
unbound
company admin/portability item event plus audit record
catalog PRP status: audit_pending
upsert_case optional_agent_tool cases:write standard
skill_test
company_write required no mock_extension:cases.upsert scenario
scenario_mock
unbound
semantic-operation item event plus company-entity state diff and audit record
catalog PRP status: audit_pending
write_document always_agent_tool none standard
planning
skill_test
task_write required no semantic_command:write_document scenario + live
live_codex
unbound
semantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events
catalog PRP status: audit_pending

Axis 3: the 16 behavioral eval groups

Behavior groups describe expected outcomes and trajectories. They do not grant tools and they do not define PRP event types. The matrix is generated from the behavior-group source plus the checked-in eval traceability manifest; scenario counts and membership cannot be hand-edited here.

Coverage matrix

Group Owner Semantic operations Control-plane operations Real Paperclip surface Mock state Scenarios PRP evidence Gap / disposition
hb — Heartbeat control plane + always tools get_task_context select_work
enforce_budget
agent identity, inbox-lite, heartbeat context, budget and active-issue services company
actor
wake
task
budget
run
5 runner/session/run context plus bounded task-context tool results Production semantic binding is unbound; identity and work selection remain injected/control-plane-owned.
co — Checkout control plane none checkout_task POST /api/issues/:id/checkout and execution-lock services task
actor
run
idempotency
fault
6 run preparation and issue-status decision evidence with checkout receipt Intentionally no model tool; the production checkout receipt still needs the additive semantic-receipt envelope.
st — Status always tools + control-plane arbitration answer_status_question
finish_task
block_task
request_review
reconcile_run
append_audit_record
issue PATCH, review/liveness policy, and native finalization arbitration task
comments
interactions
blockers
audit
run
8 semantic operation receipt, work assessment, issue-status decision, and terminal causality Production semantic binding and additive typed operation/conflict receipts remain unimplemented.
cm — Comments always tools get_task_history
report_progress
append_audit_record issue comment list/get/create routes task
comments
actor
idempotency
audit
6 bounded read result or idempotent comment-write receipt plus audit reference Active-task binding is unbound; cross-task comment mutation is deliberately outside V1.
se — Search optional discovery tools search_tasks
list_agents
get_agent
list_projects
list_goals
none company issue search and agent/project/goal list/get routes company
task
actor
project
goal
4 bounded redacted read projections through tool-result item events Project and goal operations are scenario-only; every real service binding is unbound.
su — Subtasks optional delegation tools create_task route_wake company issue create, child issue, assignment, and wake services company
task
actor
blockers
wake
audit
4 company/task state diff, audit reference, and continuation wake evidence create_task is production-bound to ordinary active-issue child creation with assignment, dependency-ready wake, company checks, child limits, and durable source-scoped idempotency.
bl — Blockers always/optional tools + control plane block_task
set_dependencies
schedule_blocker_wake
route_wake
issue relations, blocker projection, liveness validation, and blocker wake services task
blockers
wake
actor
audit
fault
5 dependency diff, block receipt, attention routing, and issue-status decision set_dependencies is production-bound for the active issue; block_task remains unbound, and cancelled-blocker receipts still need typed additive evidence.
dp — Documents and plans always tools; restore optional; destructive lifecycle control-plane-only list_documents
read_document
list_document_revisions
write_document
append_audit_record issue document list/read/upsert/revision/restore/lock/unlock/delete routes task
documents
interactions
idempotency
audit
fault
3 bounded reads and revision-safe write/conflict/denial receipts with revision lineage restore_document_revision is an approved optional-tool gap; lock/unlock/delete are intentionally control-plane-only.
ix — Interactions always tools + addressed resolver request_human_input route_wake issue-thread interaction create/respond/accept/reject/withdraw services task
documents
interactions
wake
actor
idempotency
9 interaction proposal/materialization/response/delivery and attention events Production binding and semantic resolution receipts are unbound.
ap — Approvals optional governance tools + governed approver list_approvals
get_approval
get_approval_context
request_approval
decide_approval
comment_on_approval
route_wake
append_audit_record
company approval, decision, issue-link, comment, and governed-action services company
task
approvals
actor
wake
audit
idempotency
6 governed semantic receipts, audit references, and attention/continuation linkage Production binding and additive governed-action receipts are unbound; board-only authority stays outside grants.
ar — Artifacts always tools + artifact/work-product services register_deliverable append_audit_record attachment upload and issue work-product routes task
artifacts
workProducts
workspace
audit
idempotency
4 artifact/work-product reference and durable inspectability receipt; never binary bytes Production upload/register composite and additive durable-reference receipt are unbound.
er — Errors and critical rules runner/control plane + optional workspace/wake tools get_workspace_runtime
control_workspace_service
schedule_wake
inspect_operation_result
release_task
enforce_budget
persist_run
replay_run
reconcile_run
workspace runtime, monitor/recovery, budget, run persistence/replay, release, and terminal services workspace
budget
run
wake
audit
idempotency
fault
9 runtime/workspace/attention/run lifecycle, typed denials, replay facts, and terminal causality Budget stop reasons and semantic denial/conflict receipts require additive v1 envelopes; inspect_operation_result remains scenario-only.
rf — Reference files optional domain tools + test-only escape hatch list_cases
upsert_case
list_routines
manage_routine
list_company_skills
sync_company_skills
list_secret_metadata
read_secret_value
export_company
administer_company
generic_api_request
append_audit_record case, routine, company-skill, secret, portability, and administration services company
cases
routines
skills
secrets
audit
fault
22 bounded domain projections, redacted broker receipts, company diffs, and audit references These operations are scenario-only except generic_api_request, which is test-only; broad administer_company is deferred and cannot claim product coverage.
mh — Multi-hop composed semantic operations + control-plane continuation create_task
set_dependencies
request_human_input
request_approval
register_deliverable
route_wake
reconcile_run
delegation, dependency, interaction, approval, artifact, and terminal orchestration services task
blockers
interactions
approvals
artifacts
wake
run
audit
4 correlated operation receipts, state diffs, attention hops, work assessment, status decision, and terminal outcome No generic transaction tool is allowed; shared mock/real conformance must prove each composed effect.
rs — Restraint and no-call policy/exposure layer answer_status_question
read_secret_value
generic_api_request
enforce_budget task-mode, secret-broker, test-scope, pause, and budget policy checks actor
task
budget
secrets
audit
fault
3 absence of forbidden effects plus typed policy denial/redaction receipts when a call is attempted Typed redaction/authorization receipts need additive v1 evidence; generic_api_request is never a product fallback.
wk — Wake situations control plane + always context/history tools get_task_context
get_task_history
schedule_wake
select_work
route_wake
wakeup requests, heartbeat context, comment/interaction/approval/blocker wake routing, and scheduled wake services wake
task
comments
interactions
approvals
blockers
run
8 attention request routing/resolution plus resumed session/run causality Production scheduling binding is unbound; control-plane routing remains non-callable.

Behavior group hb: Heartbeat

5 scenarios (legacy group 1):

  • hb-context-01 — Prefer the compact heartbeat-context route before replaying the thread
  • hb-inbox-lite-01 — Normal heartbeat starts from the compact inbox, not a raw issue listing
  • hb-pick-priority-01 — Pick-work priority prefers in_progress and skips blocked work
  • hb-scoped-wake-01 — Scoped wake payload skips identity/inbox and goes straight to checkout
  • hb-wake-comment-01 — Comment wake fetches the triggering comment first, then responds on the issue

Behavior group co: Checkout

6 scenarios (legacy group 2):

Behavior group st: Status

8 scenarios (legacy group 3):

  • st-backlog-park-01 — Postponed work is parked in backlog, not cancelled or closed
  • st-blocked-owner-01 — Blocking on another issue sets status blocked plus first-class blockedByIssueIds
  • st-crossteam-cancel-01 — Cross-team tasks are never cancelled — reassign to the manager instead
  • st-done-comment-01 — Closing a task is checkout, then PATCH done with an explanatory comment
  • st-env-blocked-notdone-01 — Impossible deliverable (absent mount, read-only prefix, no route) ends blocked with a named owner — never done or in_review
  • st-ephemeral-verify-01 — Artifacts produced outside the synced workspace — the disposition PATCH must carry a persistence caveat (or relocation), never a bare done
  • st-handback-review-01 — Board user asking for the task back gets reassigned + in_review, not done
  • st-unverified-toolchain-01 — Unrunnable mandated verification (certified toolchain unobtainable) ends blocked stating what could not be verified — no unqualified completion claim

Behavior group cm: Comments

6 scenarios (legacy group 4):

4 scenarios (legacy group 5):

Behavior group su: Subtasks

4 scenarios (legacy group 6):

Behavior group bl: Blockers

5 scenarios (legacy group 7):

Behavior group dp: Documents and plans

3 scenarios (legacy group 8):

Behavior group ix: Interactions

9 scenarios (legacy group 9):

  • ix-checkbox-01 — Subset selection from a known list is a checkbox confirmation
  • ix-checkbox-result-01 — Checkbox continuation wake acts on result.selectedOptionIds only
  • ix-confirmation-plan-01 — Plan sign-off is a request_confirmation bound to the latest revision, then in_review
  • ix-continuation-01 — request_confirmation sets a wake continuation policy when work must resume
  • ix-questions-01 — A short typed form of questions is ask_user_questions, not a comment
  • ix-stale-target-01 — A stale_target expiry means rebuild against the latest revision, fresh interaction
  • ix-suggest-tasks-01 — Proposing tasks for the board to accept uses suggest_tasks, not direct creation
  • ix-superseded-comment-01 — A superseded_by_comment expiry means address the comment, then a new interaction
  • ix-verdicts-01 — Per-item approve/reject decisions use request_item_verdicts with reasons on reject

Behavior group ap: Approvals

6 scenarios (legacy group 10):

  • ap-approval-deny-01 — A denied approval leaves the issue open with an explanatory comment
  • ap-approval-wake-01 — Approval wake reviews the approval, its issues, and closes what it resolves
  • ap-board-approval-01 — Spend needs a request_board_approval linked to the issue, then a waiting posture
  • ap-mcp-expiry-01 — An expired MCP approval means one fresh idempotent re-call, then in_review again
  • ap-mcp-gate-01 — Pending MCP tool approval means in_review posture, no retry, no done
  • ap-mcp-pathmissing-01 — approval_path_missing means stop and reroute, not retry loops or fake dispositions

Behavior group ar: Artifacts

4 scenarios (legacy group 11):

Behavior group er: Errors and critical rules

9 scenarios (legacy group 12):

Behavior group rf: Reference files

22 scenarios (legacy group 13):

Behavior group mh: Multi-hop

4 scenarios (legacy group 14):

Behavior group rs: Restraint and no-call

3 scenarios (legacy group 15):

Behavior group wk: Wake situations

8 scenarios (legacy group 16):

  • wk-ask-mode-01 — Ask-mode wake is answer-only — comment, no state or document writes
  • wk-plan-accepted-01 — Accepted-plan continuation creates subtasks, never re-plans or re-asks
  • wk-plan-annotation-01 — Plan-annotation wake revises the document and answers annotations via nested routes
  • wk-plan-directive-01 — Planning-mode wake produces plan PUT, revision-bound confirmation, then in_review
  • wk-plan-directive-02 — Layer probe — planning directive via wake prose only (no task-context markdown)
  • wk-recovery-processlost-01 — Process-lost recovery verifies durable progress first and never redoes posted work
  • wk-resume-delta-01 — Layer probe — resume delta with condensed contract only preserves close discipline
  • wk-skilltest-mode-01 — Skill-test mode writes the structured result to the output document, scoped to this issue

Authorization and execution invariants

Company and actor authorization

  • Every entity read and write is resolved inside the authenticated actor's company; cross-company identifiers fail without disclosing protected facts.
  • Board actors use active membership and role permissions. Agent writes require a company-scoped run JWT and X-Paperclip-Run-Id; active-task tools cannot accept a caller-selected company or arbitrary task.
  • Optional tools are omitted unless every required claim, role, and task-mode condition is satisfied. A grant never bypasses approval, budget, pause, execution-lock, interaction-owner, or other governed-action checks.

Task modes and exposure

  • standard permits the full eligible catalog subject to claims and role checks.
  • ask is answer-only: investigation reads are eligible, while implementation, document mutation, delegation, terminal, and governed writes are absent.
  • planning permits plan/document/progress/input operations but not delegation or terminal completion; accepting a plan transitions the issue into a fresh standard execution continuation for the exact accepted revision.
  • skill_test exposes only scenario-admitted tools. generic_api_request additionally requires the explicit test grant and allowlist and never supplies product coverage.

Redaction

  • Tool descriptors declare observable redaction. Secret plaintext may reach only the model-use capsule authorized for a bound secret; it never reaches PRP, logs, artifacts, errors, audit details, or serialized tool results.
  • Authorization denials reveal stable codes, missing public claims, and remediation only; they do not include protected entity state, credentials, headers, cookies, or another company's policy.
  • Artifacts carry durable references, hashes, content type, and size on the wire; binary payloads remain in storage transports.

Idempotency and replay

  • Every mutation whose descriptor says required must carry a stable idempotency key. Retrying the same key and equivalent input returns the prior receipt without duplicating side effects.
  • Reusing a key with different input is an idempotency conflict. Optimistic document writes additionally bind baseRevisionId and return the current revision on conflict.
  • PRP source event IDs are idempotent, source sequence gaps are evidence, and replay is side-effect free. Replaying a trace cannot repeat control-plane writes.

Side effects and terminal ownership

  • read operations return bounded projections and do not mutate control-plane state.
  • task_write and company_write operations emit normalized state diffs and immutable audit references after production authorization succeeds.
  • governance, workspace_control, secret_read, and admin operations retain their production-specific gates; catalog exposure is necessary but never sufficient authority.
  • Terminal semantic operations propose intent. The control plane owns checkout, release, budget enforcement, persisted run evidence, replay, wake routing, audit append, and final status arbitration.

Issue-document lifecycle

Issue documents are active-task working records identified by (activeIssueId, key). Always-present document tools do not accept a company id or arbitrary issue id.

Cross-task document reads require a future explicit optional operation and grant; case documents and future company knowledge are separate resource types.

Locked-document fallback may create a deterministic new key only when lockedDocumentStrategy is create_new_document; the source stays locked and the returned receipt names the redirected key.

Action Semantic operation Placement Contract Status
create write_document always_agent_tool baseRevisionId is null; create revision 1 and return document/revision/audit identity. supported contract; production binding unbound
read list_documents / read_document always_agent_tool List metadata or read the current active-task document by stable key. supported contract; production binding unbound
update write_document always_agent_tool Require the exact latest baseRevisionId and an idempotency key; append an immutable revision. supported contract; production binding unbound
revisions list_document_revisions always_agent_tool Return bounded immutable revision history newest first. supported contract; production binding unbound
restore restore_document_revision optional_agent_tool (documents:restore) Validate same-document lineage, reject locked documents, append a new revision, and return source/new revision identity; restoring current is a no-change duplicate. approved catalog gap
lock none control_plane_owned Board/administrative lifecycle action; agent writes receive document_locked or use explicit create-new fallback. intentionally non-callable
unlock none control_plane_owned Board/administrative lifecycle action; never inferred from a failed write. intentionally non-callable
delete none control_plane_owned Destructive audited action; pending targets become stale and the runner never recreates them automatically. intentionally non-callable

All document writes produce a normalized receipt. Success includes the document key/id, prior and new revision identity, idempotency key, and audit reference. Missing/stale bases return base_revision_required or stale_base_revision with the current revision. Authorization, cross-company, and lock failures use stable redacted denial codes. The runner never overwrites after a conflict without a fresh read and explicit reconciled write.

Provenance, generation, and drift gates

The machine-readable authority for this document's decisions is spec/operation-groups/source.json. The generator joins that source to these independently versioned authorities and rejects disagreement:

  • the package-exported CAPABILITY_CANONICAL_CATALOG in src/catalog/ — the 41-operation authority, including enriched mock/redaction and divergence facts.
  • src/tools/capability-semantic-tool-types.ts — 10 control-plane-owned operation IDs.
  • protocol/schemas/event.schema.json and command.schema.json — PRP event/command members and families.
  • spec/capability/eval-traceability.yaml — all 16 groups and 106 scenario identities/source anchors.
  • spec/capability/source-contract.json and mcp-tool-map.yaml — legacy MCP placement and fold authority.
  • generated/capability/semantic-tool-contracts.json — generated provider contract; it must equal the live catalog exactly.
  • src/catalog/index.ts and src/index.ts — package export path for the canonical reconciliation authority.

Regenerate and check reproducibly:

pnpm --filter @paperclipai/paperclip-runner exec tsx scripts/generate-operation-groups.ts
pnpm --filter @paperclipai/paperclip-runner exec tsx scripts/generate-operation-groups.ts --check
pnpm --filter @paperclipai/paperclip-runner exec vitest run src/catalog/operation-groups-doc.test.ts src/catalog/reconciliation.test.ts src/catalog/catalog-docs.test.ts

The --check path fails on catalog membership, optional-group coverage, control-plane coverage, PRP schema families/counts, behavior/scenario membership, legacy alias folds, source-contract targets, generated live contracts, package exports, or byte-level Markdown drift. Generation is offline and uses only checked-in inputs.

Current responsibility-based paths are normative. Numbered phase-* or milestone paths are historical and must not be reintroduced.

Reconciliation appendix

Catalog split and deliberate replacement

  • Scenario/eval catalog: 37 operations.
  • Live dispatcher catalog: 28 operations.
  • Shared: 24; union/canonical authority: 41.
  • Scenario-only: administer_company, export_company, inspect_operation_result, list_cases, list_company_skills, list_goals, list_projects, list_routines, list_secret_metadata, manage_routine, read_secret_value, sync_company_skills, upsert_case.
  • Live-only: get_agent, get_approval, get_approval_context, schedule_wake.
  • The generated provider contract contains exactly the live catalog; the canonical union remains the migration authority until all scenario-only operations are either implemented, deferred, or removed by an explicit reconciliation decision.
  • generic_api_request stays exported only for controlled tests and cannot be cited as real-surface, mock-parity, or PRP product coverage.

Legacy MCP aliases

The MCP inventory is a compatibility index, not a third product catalog. Every alias folds into an eval scenario and its source-contract semantic target must resolve to a canonical operation, a control-plane operation, an approved composite/fold, or a named gap.

MCP alias Source placement / target Reconciled target Eval evidence
paperclipMe control_plane_owned / injected_actor_context control_plane → select_work
Identity is launch context, not a model tool.
hb-inbox-lite-01
paperclipInboxLite control_plane_owned / runner_work_selection control_plane → select_work
Inbox selection belongs to the control plane.
hb-inbox-lite-01
paperclipListAgents optional_agent_tool / list_agents list_agents rf-api-mgr-heartbeat-01
paperclipListSkills optional_agent_tool / list_company_skills list_company_skills rf-cskill-audit-01
paperclipGetAgent optional_agent_tool / get_agent get_agent rf-api-mgr-heartbeat-01
paperclipListIssues optional_agent_tool / search_tasks search_tasks se-q-filters-01
paperclipGetIssue always_agent_tool / get_task_context get_task_context se-get-issue-01
paperclipGetHeartbeatContext always_agent_tool / get_task_context get_task_context hb-context-01
paperclipListComments always_agent_tool / get_task_history get_task_history se-get-issue-01
paperclipGetComment always_agent_tool / get_task_history get_task_history hb-wake-comment-01
paperclipListIssueApprovals always_agent_tool / get_task_context get_task_context ap-board-approval-01
paperclipListDocuments always_agent_tool / list_documents list_documents dp-base-revision-01
paperclipGetDocument always_agent_tool / read_document read_document dp-base-revision-01
paperclipListDocumentRevisions always_agent_tool / list_document_revisions list_document_revisions dp-base-revision-01
paperclipListProjects optional_agent_tool / list_projects list_projects rf-wf-project-setup-01
paperclipGetProject optional_agent_tool / get_project folded → list_projects
The bounded discovery operation owns project list/get projection in V1.
rf-wf-project-setup-01
paperclipGetIssueWorkspaceRuntime optional_agent_tool / get_workspace_runtime get_workspace_runtime rf-iws-start-url-01
paperclipControlIssueWorkspaceServices optional_agent_tool / control_workspace_service control_workspace_service rf-iws-start-url-01
paperclipWaitForIssueWorkspaceService optional_agent_tool / wait_for_workspace_service folded → control_workspace_service
Wait is a bounded action of workspace service control.
rf-iws-target-restart-01
paperclipListGoals optional_agent_tool / list_goals list_goals su-parent-goal-01
paperclipGetGoal optional_agent_tool / get_goal folded → list_goals
The bounded discovery operation owns goal list/get projection in V1.
su-parent-goal-01
paperclipListApprovals optional_agent_tool / list_approvals list_approvals ap-approval-wake-01
paperclipCreateApproval optional_agent_tool / request_approval request_approval ap-board-approval-01
paperclipGetApproval optional_agent_tool / get_approval get_approval ap-approval-wake-01
paperclipGetApprovalIssues optional_agent_tool / get_approval_context get_approval_context ap-approval-wake-01
paperclipListApprovalComments optional_agent_tool / get_approval_context get_approval_context ap-approval-deny-01
paperclipCreateIssue optional_agent_tool / create_task create_task su-parent-goal-01
paperclipUpdateIssue optional_agent_tool / semantic_task_disposition composite → answer_status_question, finish_task, block_task, request_review
Generic issue PATCH is replaced by intent-specific terminal/status operations.
st-done-comment-01
paperclipCheckoutIssue control_plane_owned / atomic_checkout control_plane → checkout_task
Checkout is an atomic control-plane transaction.
co-body-contract-01
paperclipReleaseIssue control_plane_owned / runtime_release control_plane → release_task
Release is runner/control-plane cleanup.
er-release-01
paperclipAddComment always_agent_tool / report_progress report_progress cm-multiline-01
paperclipSuggestTasks always_agent_tool / request_human_input request_human_input ix-suggest-tasks-01
paperclipAskUserQuestions always_agent_tool / request_human_input request_human_input ix-questions-01
paperclipRequestConfirmation always_agent_tool / request_human_input request_human_input ix-confirmation-plan-01
paperclipRequestCheckboxConfirmation always_agent_tool / request_human_input request_human_input ix-checkbox-01
paperclipUpsertIssueDocument always_agent_tool / write_document write_document dp-plan-doc-01
paperclipRestoreIssueDocumentRevision optional_agent_tool / restore_document_revision known_gap → named gap
Approved optional documents:restore operation is not yet in the canonical catalog.
dp-base-revision-01
paperclipLinkIssueApproval optional_agent_tool / link_approval folded → request_approval, get_task_context
Issue linkage is part of approval request/context composites.
ap-board-approval-01
paperclipUnlinkIssueApproval optional_agent_tool / unlink_approval control_plane → append_audit_record
No standalone agent unlink tool is approved in V1.
ap-board-approval-01
paperclipApprovalDecision optional_agent_tool / decide_approval decide_approval ap-approval-wake-01
paperclipAddApprovalComment optional_agent_tool / comment_on_approval comment_on_approval ap-approval-deny-01
paperclipApiRequest optional_agent_tool / test_only_api_escape_hatch folded → generic_api_request
Compatibility alias for the test-only escape hatch.
rf-api-404-report-01

PRP expressiveness boundary

PRP v1 already represents runner/session/turn/item lifecycle, replay identity, capability negotiation, request/interaction routing, result negotiation, attention, work assessment, issue-status decision, and terminal causality. Provider-neutral semantic-operation receipts (including authorization/redaction/idempotency/conflict facts) and explicit budget-stop reasons are additive v1 work. Destructive administration, cross-company actions, board-only governance, document lock/unlock/delete, checkout selection, audit persistence, and final arbitration remain control-plane-local and do not become model tools or generic PRP commands.