paperclip/packages/paperclip-runner/docs
Dotta 83987210d6
fix(runner): align direct eval provider setup with qualified runtimes (#12945)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner direct live evals test semantic tools against a mock control
plane.
> - The first complete AWS campaign exercised 358 cells.
> - It exposed setup differences from the working full-stack harness.
> - This pull request corrects those direct-harness differences.
> - It preserves production permission defaults and the full-stack
workflow.

## Linked Issues or Issue Description

**What happened?**

Native Codex cells could not find a global codex executable. ACPX denied
unattended tool requests and lost valid provider usage receipts.
AgentCore hit a 30-second facade timeout while its worker allows 120
seconds for delivery. Several models reported a native run result
without updating the separate mock task state.

**What did you expect?**

The direct harness should use the pinned executable, explicit test
permissions, and a timeout compatible with the provider delivery
contract. Its instructions should explain which operation changes mock
task state.

**Steps to reproduce**

Run the full Runner Direct Live Protocol Evals workflow. Baseline
campaign:
https://github.com/paperclipai/paperclip/actions/runs/34059009921.

**Paperclip version**

Master at b55ce03e86. Related public
change: #12932. No duplicate fix was found.

## What Changed

- Resolve native Codex from the pinned Codex ACP dependency, as the
full-stack launcher does.
- Select explicit unattended ACPX permissions only in the isolated
direct eval harness.
- Add a configurable bounded turn-admission wait. AgentCore direct evals
use up to 125 seconds, capped by their turn budget. Other callers retain
the current default waits.
- Explain mock task-state operations separately from native run-result
reporting. Keep all scoring assertions unchanged.
- Bind the direct turn's scope to the current request, rather than stale
shared fixture notes. A provider's end-of-turn result does not authorize
an unrequested mock task completion or extra completion comment.
- Normalize the qualified Claude/Codex usage semantics without
double-counting reasoning or inventing unknown billable categories.
Forward ACPX's persisted terminal prompt-response usage for the exact
current turn; reject stale or ambiguous receipts.
- Preserve the known numeric ACPX token-counter aliases through durable
redaction. Continue redacting strings and credential-shaped values. A
regression test reproduces the previously redacted usage before
normalization.
- Add regression tests and operating documentation.

## Verification

- Passed Runner TypeScript build.
- Passed 88 focused tests across the eval request contract, provider
setup, and Runner transport.
- Passed 93 focused ACPX adapter, sidecar lifecycle, and
usage-accounting tests, including persisted receipt identity and
missing-field regressions.
- Passed 97 additional runtime-host, in-process ACPX driver, OpenCode
proxy, and MCP bridge tests.
- Reproduced the ACPX counter-redaction failure, then passed 36 Rust
decoder/normalizer and durable-redaction tests after the fix.
Credential-shaped strings and objects remain redacted.
- Native Codex get-task-context passed a real provider probe with no
global CLI dependency.
- Local ACPX probes reached a platform startup rejection on macOS;
qualification therefore used the intended AWS Linux fleet. Complete
campaign 34060573948 reached 341/358 passing with zero infrastructure
failures (up from 250/358 and 100 infrastructure failures). ACPX Claude
reached 34/35, ACPX Codex 7/8, native Codex 34/35, and AgentCore 35/35.
Remaining failures were retained in the canonical report:
https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-34060573948-1/.
- Final complete campaign 34062394019 exercises the current-request
scope clarification with all 358 cells; results pending.
- Passed git diff --check.
- Repo-wide typecheck, build, and tests are delegated to PR CI. They
were not repeated on this machine.

## Risks

- ACPX approval is scoped to the operator-requested isolated fixture
harness. Production defaults do not change.
- AgentCore admission can wait longer, but remains within the whole-turn
budget.
- The mock fixture has no production task lifecycle service. Explicit
task-state instructions describe that boundary; they do not relax
scoring.
- This change does not modify the browser full-stack workflow or its
test fixtures.
- ACPX receipt corrections apply to all native ACPX sessions. They rely
on the pinned qualified server contracts; unknown or ambiguous receipts
stay unknown, and a final accounting-read error never overwrites the
provider's terminal result.

## Model Used

- OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, tool use,
and code execution. The session does not expose an exact context-window
limit.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-06 17:21:12 -05:00
..
adr feat(runner): add offline evaluation tooling (#12653) 2026-09-01 05:19:47 -05:00
design feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
research feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
tutorials feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
adding-a-harness.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
architecture.md feat(runner): add offline evaluation tooling (#12653) 2026-09-01 05:19:47 -05:00
capability-authorization-and-exposure.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-clean-room-chat.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-contract.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-disposition.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-eval-conformance.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-eval-slice.md feat(runner): add offline evaluation tooling (#12653) 2026-09-01 05:19:47 -05:00
capability-execution-modes.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-future-binding-boundary.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-issue-thread-ui.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-live-runnerd-codex.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-mock-control-plane-port.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-scenario-explorer.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-semantic-catalog.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-semantic-tools.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
capability-verification-commands.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
codex-driver.md fix(runner): restore local session and task integrity (#12721) 2026-09-02 16:11:26 -05:00
durable-recovery.md fix(runner): recover native sessions across restarts (#12845) 2026-09-04 15:03:53 -05:00
evals-integration.md feat(runner): add offline evaluation tooling (#12653) 2026-09-01 05:19:47 -05:00
index.md feat(runner): add offline evaluation tooling (#12653) 2026-09-01 05:19:47 -05:00
live-console-protocol-server.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
live-console.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
local-runner.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
protocol-compatibility.md feat(runner): add secure remote transport (#12639) 2026-09-01 02:19:11 -05:00
runner-protocol-live-evals.md fix(runner): align direct eval provider setup with qualified runtimes (#12945) 2026-09-06 17:21:12 -05:00
runner-workflow-evals.md feat(runner): restore direct live eval campaigns and reports (#12909) 2026-09-05 20:18:11 -05:00
scenario-chat.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
sdk.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00
standalone-thin-paperclip-adapter.md feat(runner): add SDK and developer tooling (#12608) 2026-08-31 21:33:11 -05:00