Commit Graph

69 Commits

Author SHA1 Message Date
Dotta f57d1e3711 ci(runner-e2e): reuse AWS Chrome in paid cells 2026-09-04 07:41:18 -05:00
Dotta c2d33c8b99 test(runner-e2e): bypass stale legacy OpenCode catalog 2026-09-04 07:41:18 -05:00
Dotta d13b9ede00 test(runner-e2e): retry exact DeepSeek terminal variance 2026-09-04 02:56:20 -05:00
Dotta fca39344b1 ci(runner): install only Chromium headless shell 2026-09-04 02:26:19 -05:00
Dotta 9b231a37a3 ci(runner): preserve evidence across failed-job reruns 2026-09-04 01:58:05 -05:00
Dotta f90e475c97 test(runner): make legacy question fixture idempotent 2026-09-04 00:47:58 -05:00
Dotta a2a623464f fix(runner-e2e): isolate provider cache after browser launch 2026-09-04 00:07:13 -05:00
Dotta 556785beab fix(opencode): refresh isolated model cache 2026-09-03 23:45:37 -05:00
Dotta 8961aee0de ci(runner): skip Daytona setup for local cells 2026-09-03 22:35:46 -05:00
Dotta b8d84471a3 fix(runner): preserve remote continuation state scope 2026-09-03 22:09:10 -05:00
Dotta 49910c1c24 Merge PR fast-path activation from master
* commit '89bf6a33c23d12c592c2c45b07ac36a362d418e6':
  ci: activate Node-first pnpm setup for PRs (#12810)
  feat(apps): unify permissions and action testing (#12802)
2026-09-03 21:29:36 -05:00
Dotta 89bf6a33c2
ci: activate Node-first pnpm setup for PRs (#12810)
## Thinking Path

> - Paperclip validates every change through an immutable reusable PR
workflow.
> - That caller still pinned a revision that ran pnpm setup before Node
setup.
> - The implementation fix in #12808 is therefore present on master but
inactive for ordinary PR CI.
> - Advancing the immutable caller pin activates the already tested
Node-first workflow.
> - A focused contract prevents the caller from silently returning to
the old revision.
> - The benefit is a faster PR feedback loop without changing product
code or secret boundaries.

## Linked Issues or Issue Description

Refs #12808

## What Changed

- Pin ordinary PR CI to trusted workflow revision
`a0a78ee60946a5f79f85b2bd0584fc766fae43bb`.
- Assert that the reusable workflow call is canonical, unique, and
SHA-pinned to that audited revision.

## Verification

- Focused workflow security test: 8/8.
- Prettier passed.
- Actionlint passed.
- `git diff --check` passed.

## Risks

Low risk. The change only advances an immutable reusable-workflow pin to
a revision whose full ordinary CI and security checks passed. Product
code and credentials are unchanged.

## Model Used

OpenAI Codex GPT-5 with agentic reasoning and repository tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used
- [x] I have linked the related public PR
- [x] I have not referenced internal issue links
- [x] My branch name describes the change
- [x] I have run focused tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered and documented risks
2026-09-03 21:28:49 -05:00
Dotta 95aa8c514e Merge master into runner paid matrix integrity
* commit 'a0a78ee60946a5f79f85b2bd0584fc766fae43bb':
  ci: bootstrap Node before pnpm setup (#12808)
  ci(runner): stamp paid target provenance (#12805)
  fix: guard listComments against non-UUID afterCommentId to prevent 500 errors (#8695)
  test(plugin-worker): remove the wall-clock race from the duplex buffered-replay tests (#12799)
  ci(runner): skip bootstrap registry telemetry (#12797)
  fix(ui): honor PAPERCLIP_HIDDEN_SETTINGS in the production switcher menu (#12788)
  chore(deps): bump motion from 12.43.0 to 13.1.1 (#12255)
  chore(deps): bump dompurify from 3.4.13 to 3.4.14 (#12266)
  chore(deps-dev): bump @types/react-dom from 19.2.4 to 19.2.5 (#12253)
  ci(runner): inspect Daytona image metadata remotely (#12795)
  fix(ui): polish core navigation and task layout (#12793)
  chore(deps): bump yjs from 13.6.29 to 13.6.32 (#12256)
  chore(deps): bump @aws-sdk/client-s3 from 3.1120.0 to 3.1122.0 (#12261)
  chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
  chore(deps): bump react-i18next from 17.0.11 to 17.0.12 (#12263)
  fix(ui): drop the "Open invite" action from the invites section (#12787)

# Conflicts:
#	tests/runner-e2e/workflow-security.test.ts
2026-09-03 21:19:07 -05:00
Dotta a0a78ee609
ci: bootstrap Node before pnpm setup (#12808)
## Thinking Path
The trusted workflows currently invoke `pnpm/action-setup@v6` before
installing the repository Node version. On hosts whose ambient Node is
older than 22.13, the action downloads standalone `@pnpm/exe`, which has
repeatedly taken several minutes. Supplying Node 24 first lets the same
pinned pnpm action use its normal Node-backed path.

## What Changed
- install Node 24 before every trusted `pnpm/action-setup` invocation
- preserve the existing pnpm-store cache setup, pinned actions,
telemetry suppression, conditions, and secret boundaries
- enforce ordering, modern Node, condition parity, and cache counts in
the workflow security contract

## Verification
- workflow security: 7/7
- Prettier
- actionlint (excluding one pre-existing SC2129 in an untouched Daytona
shell block)
- `git diff --check`

## Risks
Low. Product code, providers, paid-runner selection, credentials, and
pnpm version are unchanged. Jobs that restore pnpm cache run
`setup-node` a second time after pnpm becomes available; the first setup
is deliberately cache-free.

## Model Used
Codex (GPT-5)
2026-09-03 21:18:08 -05:00
Dotta 18ea965442
ci(runner): stamp paid target provenance (#12805)
## Thinking Path
Trusted workflow-dispatch runs execute an authorized target SHA, but
GitHub context still describes the default-branch workflow revision.
Retained paid results and artifact names were therefore labeling
target-branch executions as master. The workflow must explicitly pass
its authorized target coordinates to target code and trusted reporting.

## What Changed
- emit the canonical authorized target ref alongside the immutable
target SHA
- pass those coordinates to paid cells and the trusted report
- name shared build/provider artifacts with the target SHA rather than
workflow SHA
- add workflow-security coverage for all trusted provenance wiring

## Verification
- focused workflow-security tests: 6/6 passed
- Prettier and git diff checks passed
- run 33823252706 independently proved the pre-fix defect: functionally
green target cells were retained as master SHA 0ad180b85 instead of
feature SHA 33c7646d3

## Risks
The execution checkout and secret boundary were already pinned
correctly; this changes retained attribution and artifact labels only.
Target-side report code on PR #12769 consumes these trusted environment
values and overwrites untrusted cell metadata.

## Model Used
Codex (GPT-5)
2026-09-03 21:15:47 -05:00
Dotta 0ad180b85f
ci(runner): skip bootstrap registry telemetry (#12797)
## Thinking Path
Every trusted PR and paid-workflow job invokes the pinned pnpm setup
action. Its internal npm install is currently waiting four to seven
minutes on npm audit telemetry before any Paperclip or provider code
runs. Audit, funding, and update notifications are not integrity
controls for this action; its committed lockfile still verifies
installed package bytes.

## What Changed
- disable npm audit, funding, and update-notifier telemetry narrowly on
all seven pinned setup steps in each of the trusted PR and full-stack
workflows
- add a workflow security contract proving every setup invocation
remains covered and the overrides do not leak elsewhere

## Verification
- focused workflow security tests: 6/6 passed
- Prettier and git diff checks passed
- observed unhealthy setup: 4-7+ minutes; historical healthy setup:
about four seconds

## Risks
This skips npm vulnerability-report telemetry for the setup action
bootstrap only. Repository dependency checks, lockfile integrity,
provider-secret authorization, and target-lock verification remain
unchanged.

## Model Used
Codex (GPT-5)
2026-09-03 19:47:46 -05:00
Dotta 33c7646d3f fix(runner-e2e): record trusted target provenance 2026-09-03 19:22:00 -05:00
Dotta 97f771e2ec ci(runner): align Daytona cache exclusions 2026-09-03 19:11:39 -05:00
Dotta 03faa644fb
ci(runner): inspect Daytona image metadata remotely (#12795)
## Thinking Path
The reused Daytona image path already verifies the signed immutable
digest. It then downloads every filesystem layer only to read OCI config
fields. Buildx can retrieve the same config from that immutable digest
without pulling the layers. The assertions can therefore stay intact
while removing the expensive transfer.

## What Changed
- inspect the signed immutable Daytona image config through Buildx after
GHCR logout
- preserve digest, source revision, content ID, platform, user, and
provider-pack assertions
- extend the workflow contract test for the metadata-only path

## Verification
- Daytona image and workflow security tests: 10 passed
- Prettier and git diff checks passed
- observed full pull/prune cost: about 4m55s; metadata inspection: about
one second

## Risks
The current image has one runnable linux/amd64 platform plus its
attestation. A future genuinely multi-platform image would need explicit
linux/amd64 selection.

## Model Used
Codex (GPT-5)
2026-09-03 18:10:38 -05:00
Dotta 7fe94196ba test(runner): stage the build-once binary remotely 2026-09-03 17:45:10 -05:00
Dotta 73a146e076 fix(runner): separate remote artifact identity from launch path 2026-09-03 17:40:21 -05:00
Dotta d3c04d8932
fix(runner-e2e): prepare frozen Daytona plugin dependencies (#12791)
## Thinking Path

> - Paperclip manages AI agents and their provider runtimes.
> - The paid runner workflow installs target dependencies with lifecycle
scripts disabled.
> - The bundled Daytona plugin depends on an audited repo-local plugin
SDK link.
> - The lifecycle-safe install path did not create that link.
> - This pull request restores only the trusted Daytona preparation step
before provider secrets are exposed.
> - The benefit is a working Daytona canary without enabling dependency
lifecycle scripts.

## Linked Issues or Issue Description

**What happened?**

The Daytona paid canary stopped before lease or provider startup. The
trusted paid job disabled root lifecycle scripts, so the repo-local
plugin SDK link was absent. The plugin install returned a missing
runtime dependency error for @paperclipai/plugin-sdk.

**Expected behavior**

The trusted workflow must prepare the bundled Daytona plugin without
running untrusted dependency lifecycle scripts. The paid cell must start
only after its runtime dependencies and entrypoints pass validation.

**Steps to reproduce**

1. Dispatch the runner full-stack paid workflow for
core-compatibility.runner-acpx-claude.daytona.message-marker.
2. Let the trusted job install root dependencies with lifecycle scripts
disabled.
3. Observe the Daytona plugin installation fail before a lease or
provider process starts.

**Paperclip version or commit**

Feature head 781ac7e08c. The failed canary
is Actions run 33803959325.

**Deployment mode**

GitHub Actions paid runner validation.

**Agent adapter(s) involved**

ACPX Claude through the bundled Daytona plugin.

**Additional context**

This is a small trusted-workflow prerequisite for public PR #12769. Old
green run 33118525827 created the SDK link through root postinstall.
This change keeps lifecycle scripts disabled and restores only the
audited prerequisite.

## What Changed

- Install standalone Daytona dependencies with lifecycle scripts
disabled.
- Run the audited repo-local plugin SDK linker before provider secrets
are exposed.
- Build the bundled Daytona plugin and verify its runtime dependency
plus both entrypoints.
- Add a security regression for ordering, scope, and secret isolation.

## Verification

- Five focused workflow-security tests passed.
- Seven focused linker tests passed.
- The exact Daytona preparation command completed locally in nine
seconds.
- Prettier and diff whitespace checks passed.

## Risks

Risk is low and limited to Daytona paid cells. The setup still disables
dependency lifecycle scripts. The trusted step runs before provider
credentials enter the job. Any missing or mismatched path fails closed
before provider startup.

## Model Used

OpenAI GPT-5 Codex with repository tools and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used with version and capability
details
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in this PR with the bug template labels
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change
- [x] Focused local tests pass
- [x] I added tests for the change
- [x] I updated the relevant trusted-workflow security regression
- [x] I documented the risks above
2026-09-03 16:31:22 -05:00
Dotta 6e7f72a906 fix(runner-e2e): freeze Daytona plugin dependencies 2026-09-03 16:17:21 -05:00
Dotta 5871ebefc6 fix(runner-e2e): prepare Daytona plugin before paid cells 2026-09-03 16:10:39 -05:00
Dotta c087cf6737 test(runner): end legacy question turn before resume 2026-09-03 14:44:19 -05:00
Dotta f4686d9faa test(runner): quarantine unqualified Tencent plan cell 2026-09-03 14:41:18 -05:00
Dotta 313d6ca115
fix(runner): materialize pinned OpenCode binary (#12782)
## Thinking Path

> - Paperclip manages AI agents and their provider runtimes.
> - Paid runner validation installs target dependencies with lifecycle
scripts disabled.
> - OpenCode leaves a sentinel executable until its package lifecycle
script runs.
> - Running arbitrary lifecycle code would weaken the paid-secret
boundary.
> - This pull request materializes one exact pinned binary before
secrets are exposed.
> - The benefit is working OpenCode validation without trusting
dependency install scripts.

## Linked Issues or Issue Description

**What happened?**

Every local OpenCode paid cell stopped before provider startup because
`pnpm install --ignore-scripts` correctly retained
`opencode-ai/bin/opencode.exe` as a sentinel.

**Expected behavior**

The trusted workflow must make the exact lockfile-pinned OpenCode
executable available without running package lifecycle scripts.

**Steps to reproduce**

Run a local legacy or native OpenCode paid cell from the trusted
workflow after the target dependency install. The provider health check
reports that the OpenCode postinstall script was not run.

**Paperclip version or commit**

Default branch commit `865b4854fb44d3689f1c0ff17e3e715d52aaea73`.

## What Changed

- Materialize only `opencode-linux-x64-baseline@1.18.17` into the
matching `opencode-ai@1.18.17` package.
- Verify package identity, version, regular-file type, SHA-256 equality,
executable permissions, and runtime `--version`.
- Invoke the helper for local OpenCode and breadth cells and for remote
provider-pack assembly.
- Retain `pnpm install --ignore-scripts`.
- Add helper and trusted-workflow security regressions.

## Verification

- Helper syntax checks passed.
- Helper unit tests passed: 2/2.
- Workflow-security tests passed: 5/5.
- Prettier, actionlint, and diff whitespace checks passed.

## Risks

Risk is low and contained to paid runner setup. The helper supports only
Linux x64, fails closed on package or version drift, and runs before
provider credentials enter the job.

## Model Used

OpenAI GPT-5 Codex with repository tools and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change.
- [x] I have specified the model used.
- [x] I have checked ROADMAP.md and confirmed this does not duplicate
planned core work.
- [x] I have searched GitHub for duplicate or related PRs and found
none.
- [x] I have described the issue in this PR with the bug template
labels.
- [x] I have not referenced internal or instance-local issues.
- [x] My branch name describes the change.
- [x] Focused local tests pass.
- [x] I added tests for the change.
- [x] I updated the runner E2E security documentation.
- [x] I documented the risks above.
2026-09-03 14:12:17 -05:00
Dotta a673fca0a3 ci(runner): materialize pinned OpenCode binary 2026-09-03 13:57:53 -05:00
Dotta d238d6c01e fix(runner-e2e): reserve conflict-free local ports 2026-09-03 13:56:17 -05:00
Dotta 8952181bb3 Merge remote-tracking branch 'origin/master' into fix/runner-paid-matrix-integrity
* origin/master:
  ci(runner): build paid artifacts once per campaign (#12777)
  chore(deps): bump lucide-react from 1.32.0 to 1.38.0 (#12313)

# Conflicts:
#	.github/workflows/runner-full-stack-e2e.yml
#	tests/runner-e2e/README.md
#	tests/runner-e2e/SECURITY.md
#	tests/runner-e2e/history.test.ts
#	tests/runner-e2e/workflow-security.test.ts
2026-09-03 13:35:45 -05:00
Dotta 865b4854fb
ci(runner): build paid artifacts once per campaign (#12777)
## Thinking Path

> - Paid cells repeated the same TypeScript and Rust builds even when
one campaign selected dozens of cells.
> - The trusted workflow can compile once without provider credentials
and distribute run-scoped, digest-verified artifacts.
> - The paid cell can then disable install lifecycle scripts, verify
each artifact before extraction, and expose provider credentials only to
the final test step.
> - Local JS-backed providers also need the setup-node interpreter
permission-qualified before Rust verifies the launch artifact.

## Linked Issues or Issue Description

Run 33786122875 proved target-lock setup and catalog selection, then
failed before provider creation because trusted master did not yet
qualify the setup-node interpreter. The same workflow also rebuilt
TypeScript and Rust inside every matrix cell.

## What Changed

- Build runner TypeScript and native binaries once per campaign in a
credential-free job.
- Build the remote provider pack once only when selected Daytona cells
require it.
- Upload run-scoped bundles with SHA-256 manifests and verify before
extraction in each paid cell.
- Remove repeated TypeScript, provider-pack, and Rust builds from paid
cells.
- Qualify the local provider Node interpreter before verified launch.
- Propagate the resolved target lockfile through all five target-code
jobs.
- Keep local-only selection off Daytona and exclude Xiaomi from the
67-cell catalog.

## Risks

A shared build artifact could fan out a bad payload to many cells. The
producing jobs receive no provider credentials, use the exact authorized
target SHA and resolved lockfile, and publish run-scoped artifacts.
Every consuming job verifies SHA-256 before extraction. Paid dependency
setup keeps lifecycle scripts disabled and provider credentials remain
scoped to the final test step.

## Verification

- Focused runner workflow-security, catalog, and Daytona-image tests:
25/25 passed.
- Prettier passed.
- Actionlint passed with only the two pre-existing SC2129 style notices
ignored.
- Git diff check passed.

## Model Used

OpenAI Codex, GPT-5.

## Checklist

- [x] Build jobs are credential-free.
- [x] Paid installs disable lifecycle scripts.
- [x] Artifacts are run-scoped and digest-verified before extraction.
- [x] Trusted report and history jobs remain isolated from target
artifacts.
- [x] No Daytona or Xiaomi paid run was started for this change.
2026-09-03 13:12:58 -05:00
Dotta 9e5f0e60fa Merge remote-tracking branch 'origin/master' into fix/runner-paid-matrix-integrity
* origin/master:
  ci(runner): prepare target lockfile once for paid validation (#12774)
  chore(deps): bump paperclipai/paperclip/.github/workflows/pr-trusted.yml from 39b8ee2960 to f038633bf5 (#12562)
  chore(deps): bump sharp from 0.35.3 to 0.35.4 (#12563)
  chore(deps): bump @mdxeditor/editor from 4.2.1 to 4.2.3 (#12564)
  fix(server): stop paging Sentry for supervised boot races in managed cloud (#12772)
  Secure Cloud canonical runtime identity (#12766)

# Conflicts:
#	.github/workflows/runner-full-stack-e2e.yml
#	tests/runner-e2e/README.md
#	tests/runner-e2e/SECURITY.md
#	tests/runner-e2e/workflow-security.test.ts
2026-09-03 12:42:04 -05:00
Dotta 6e50ca9d0a
ci(runner): prepare target lockfile once for paid validation (#12774)
## Thinking Path

> - The trusted target-branch runner workflow checks out PR code before
paid tests.
> - PR policy intentionally forbids manual lockfile commits.
> - Some runner changes legitimately alter pnpm patch hashes.
> - Frozen installs therefore fail before test selection.
> - Resolve one script-disabled lockfile from the authorized immutable
target SHA and distribute it by exact artifact ID and digest.
> - Keep provider credentials and trusted reporting outside this
resolution job.

## Linked Issues or Issue Description

Target-branch paid runner campaigns currently fail frozen install when a
PR changes pnpm patch content, even though ordinary PR CI regenerates
the lockfile.

## What Changed

- Added one credential-free target-lock job that resolves the authorized
immutable target SHA with lifecycle scripts disabled.
- Uploaded the resolved lockfile with its SHA-256 and restored it by
exact artifact ID before every target-code frozen install.
- Left trusted reporting and history jobs on the workflow SHA.
- Changed the disabled-AWS fallback from unavailable ubuntu-latest-m to
ubuntu-latest.

## Risks

The workflow evaluates pnpm lockfile resolution from authorized target
code. That job receives no provider credentials, disables lifecycle
scripts, rejects unrelated workspace mutations, and exposes only a
digest-verified lockfile artifact. Paid-secret jobs consume only that
lockfile after exact artifact-ID and SHA-256 validation.

## Verification

- Runner workflow-security focused tests pass.
- actionlint passes.
- Prettier and git diff checks pass.

## Model Used

OpenAI Codex, GPT-5.

## Checklist

- [x] Change is narrowly scoped to paid runner orchestration.
- [x] Target lock resolution has no provider credentials and disables
lifecycle scripts.
- [x] Downloaded artifacts are selected by exact artifact ID and
verified by SHA-256.
- [x] Trusted reporting and history jobs remain on the workflow SHA.
2026-09-03 12:39:18 -05:00
Dotta eefc36d132 ci(runner): build target branch campaign artifacts 2026-09-03 12:01:36 -05:00
Dotta 0f140a3433 Merge remote-tracking branch 'origin/master' into fix/runner-paid-matrix-integrity
* origin/master:
  chore(deps): bump better-auth from 1.7.0 to 1.7.2 (#12565)
  chore(deps-dev): bump @vitejs/plugin-react from 4.7.0 to 6.1.1 (#12566)
  chore(lockfile): refresh pnpm-lock.yaml (#12771)
  chore(deps): bump @aws-sdk/client-s3 from 3.1115.0 to 3.1120.0 (#12567)
  ci(runner): allow trusted branch targets (#12768)
  chore(deps): bump @tanstack/react-query from 5.101.4 to 5.102.8 (#12568)
  chore(deps): bump mermaid from 11.16.1 to 11.17.2 (#12569)

# Conflicts:
#	.github/workflows/runner-full-stack-e2e.yml
2026-09-03 12:01:28 -05:00
Dotta f773b2004d test(ci): accept formatted paid job dependencies 2026-09-03 11:59:37 -05:00
Dotta bf51b67285 test(runner): add isolated local provider smoke loop 2026-09-03 11:59:32 -05:00
Dotta 98c569b2df
ci(runner): allow trusted branch targets (#12768)
## Thinking Path

> - Paperclip uses paid runner tests to qualify agent execution.
> - The runner workflow controls provider secrets and AWS runner access.
> - The trusted workflow must stay on the protected default branch.
> - The code under test often exists on a branch before merge.
> - CODEOWNERS need a safe way to select that branch.
> - This pull request separates workflow authority from the code under
test.
> - The benefit is pre-merge AWS testing without target-controlled
workflow code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The manual Runner Full-Stack E2E workflow can test only the default
branch.

**Subsystem affected**

GitHub Actions and the paid runner E2E security boundary.

**Current behavior**

A CODEOWNER must merge runner changes before the trusted AWS workflow
can test them.
Selecting another branch as the workflow ref is rejected.

**Proposed behavior**

A CODEOWNER starts the workflow from `master` and supplies a
same-repository branch in `target_branch`.
The authorization job resolves the branch to one commit SHA.
Catalog, image, and paid test jobs check out that SHA after
authorization.
Report sanitization and AWS publication use the trusted workflow SHA.

**Reason and benefit**

This permits paid pre-merge qualification on AWS.
It keeps the workflow definition, report sanitizer, history publisher,
environment deployment, and runner-group permission on `master`.

**Breaking changes**

None.
The new input is optional.
An omitted input still tests the default branch.

## What Changed

- Add the optional `target_branch` workflow input.
- Resolve only a branch in `paperclipai/paperclip` to an immutable SHA.
- Pin catalog, image, paid test, and Daytona provenance to the target
SHA.
- Pin report sanitization and AWS history publication to the trusted
workflow SHA.
- Disable persisted checkout credentials in every job.
- Key cancellation by the selected target branch.
- Add policy regression coverage and operator documentation.

## Verification

- `pnpm test:e2e:runner:unit` passes with 65 tests.
- `actionlint -ignore SC2129
.github/workflows/runner-full-stack-e2e.yml` passes.
- Prettier checks pass for all changed files.
- `git diff --check` passes.

## Risks

A CODEOWNER can authorize selected branch code to receive a cell-scoped
provider credential.
This is the intended trust decision.
The workflow rejects fork refs and target-controlled workflow
definitions.
The trusted workflow SHA owns report sanitization and AWS history
publication.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5.
The exact serving snapshot and context-window size are not exposed.
The model used tool-enabled reasoning and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-03 10:38:56 -05:00
Dotta 4b7af1be41 test(runner): retire unqualified DeepSeek plan cell 2026-09-03 09:52:34 -05:00
Dotta 6b03239e16 ci(runner): reuse Daytona images for test-only changes 2026-09-03 09:49:34 -05:00
Dotta 5b71c4c49a test(runner): retry proven native plan variance 2026-09-03 09:29:44 -05:00
Dotta e0ad6d356f test(runner): retry proven breadth plan variance 2026-09-03 09:26:03 -05:00
Dotta 2311f08422 test(runner): align terminal fixture ordering 2026-09-03 09:02:08 -05:00
Dotta 1b74561fea
ci(runner): route paid matrix to AWS fleet (#12765)
## Thinking Path

> - Paperclip manages AI agents that perform work.
> - The paid runner matrix verifies complete runner behavior with real
providers.
> - Each matrix job currently repeats work on GitHub-hosted runners.
> - Paperclip has an ephemeral AWS runner fleet for trusted workflows.
> - The paid workflow needs a reviewed and fail-closed route to that
fleet.
> - This pull request adds that route and keeps the existing hosted
runner as the disabled-state fallback.
> - The benefit is faster paid campaigns with the same actor,
environment, and secret boundaries.

## Linked Issues or Issue Description

**What happened?**

The Runner Full-Stack E2E workflow always uses `ubuntu-latest-m`. It
limits the matrix to 57 parallel jobs. The repository AWS fleet can run
100 ephemeral jobs, but the paid workflow cannot select it.

**Expected behavior**

An explicit repository flag must select the reviewed AWS fleet label. A
missing or invalid flag must keep the existing hosted runner. The
workflow must authorize the stable actor identity before it routes any
paid job.

**Steps to reproduce**

1. Dispatch the Runner Full-Stack E2E workflow from `master`.
2. Inspect a paid matrix job.
3. Observe that the job requests `ubuntu-latest-m` even when the AWS
fleet should be used.

**Paperclip version or commit**

`da0947d3582ac7779d6bf11851c9938eca6c5c8c`

**Deployment mode**

GitHub Actions paid runner campaign.

## What Changed

- Add a fail-closed `RUNNER_E2E_AWS_ENABLED` switch.
- Select only the reviewed AWS fleet label or the existing hosted label.
- Permit up to 100 parallel jobs in AWS mode.
- Keep the hosted-runner limit at 57.
- Reauthorize paid execution before checkout and provider access.
- Stop paid checkouts from storing GitHub credentials.
- Cancel superseded validation-ref campaigns while preserving `master`
audit runs.
- Add workflow policy checks and operator documentation.

## Verification

- `git diff --check`
- `actionlint -ignore SC2129
.github/workflows/runner-full-stack-e2e.yml`
- The organization runner group permits this workflow only from
`refs/heads/master`.
- The repository AWS switch remains disabled until this pull request is
merged and a one-cell probe succeeds.

## Risks

- A wrong fleet policy can leave jobs queued. The disabled state keeps
the existing hosted runner.
- The AWS fleet uses paid compute. The workflow validates a configured
maximum of 100 jobs.
- The runner group, actor allowlist, and paid environment remain
separate enforcement layers.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5 with agentic reasoning, repository
inspection, code editing, Git, GitHub API coordination, and static
workflow analysis. The exact deployed model identifier and
context-window size are not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-03 08:50:36 -05:00
Dotta e228846d4a test(runner): retire Xiaomi breadth cells 2026-09-03 08:49:05 -05:00
Dotta 05c7f3a857 style(runner): format paid workflow repairs 2026-09-03 08:38:33 -05:00
Dotta 23de709aa9 ci(runner): use proven hosted fallback runner 2026-09-03 08:36:56 -05:00
Dotta 5961207ce8 test(runner): import plan reset observation helper 2026-09-03 08:35:31 -05:00
Dotta f9f798f320 test(runner): allow accepted plan ACPX session resets 2026-09-03 08:34:08 -05:00
Dotta 41a96da703 ci(runner): skip reports for cancelled campaigns 2026-09-03 08:22:02 -05:00