Commit Graph

3 Commits

Author SHA1 Message Date
Dotta 8430bd897f
ci: reuse trusted cache for Daytona images (#12862)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The full-stack runner campaign checks local and Daytona runner
behavior.
> - A Daytona image content miss starts a cold multi-stage Docker build.
> - Stable dependency and agent CLI layers take most of the image build
time.
> - Development targets must not write shared cache state.
> - This pull request adds a registry cache with a default-branch write
gate.
> - It also puts volatile source inputs after stable install layers.
> - The benefit is a shorter Daytona image build without weaker secret
isolation.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the Daytona runner image stage in the full-stack E2E
workflow.

**Subsystem affected**

The GitHub Actions runner E2E workflow and its Daytona Docker image are
affected.

**Current behavior**

Each new Daytona image content ID starts with an empty BuildKit cache. A
runner source change also invalidates dependency and agent CLI install
layers because volatile inputs occur before those layers.

**Proposed behavior**

All authorized campaigns can read one GHCR BuildKit cache. Only a
campaign whose target ref is the repository default branch can update
that cache. The Dockerfile installs dependencies and agent CLIs before
it consumes volatile runner source or revision metadata.

**Reason and benefit**

The paid runner matrix spends several minutes building the image before
any selected cell can start. Cache reuse removes repeated stable setup
work and makes focused Daytona iterations faster.

**Breaking changes**

None. The immutable content tag, digest inspection, Cosign signature,
image labels, pinned base images, and provider credential boundary stay
unchanged.

## What Changed

- Read a registry-backed BuildKit cache for Daytona image content
misses.
- Export the cache only when the resolved target ref is the default
branch.
- Keep provider credentials outside the image build and cache.
- Install provider-pack dependencies before runner source is copied.
- Keep expensive agent CLI installs before source revision metadata.
- Add workflow and Docker layer-order contract checks.

## Verification

- `prettier --write .github/workflows/runner-full-stack-e2e.yml
tests/runner-e2e/daytona-image.test.ts
tests/runner-e2e/workflow-security.test.ts`
- `actionlint .github/workflows/runner-full-stack-e2e.yml`
- `git diff --check`
- I did not run a test suite or Docker image build locally. The
requested iteration policy reserves those checks for GitHub Actions.

## Risks

Low risk. BuildKit can use a cache record only when its content key
matches the build instruction and input. Development targets have
read-only cache access. The cache contains public source and build
outputs, but it does not receive provider credentials or the GitHub
token as Docker build inputs.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-05 06:31:27 -05:00
Dotta d593463ab6
perf(e2e): narrow Daytona image cache inputs (#12850)
## Thinking Path

> - Paperclip uses paid full-stack tests to verify local and Daytona
runner behavior.
> - Daytona tests reuse a content-addressed runner image when its
runtime inputs match.
> - The prior key covered the full runner package even when Docker
excluded development files.
> - Test-only and documentation changes could therefore force an
identical image rebuild.
> - This pull request aligns the Docker input closure and content-key
closure.
> - The benefit is faster paid-test iteration without unsafe image
reuse.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The Daytona paid-test workflow currently rebuilds its large runner image
after changes to runner tests, fixtures, smoke scripts, or
documentation. Those files do not enter the image and do not change its
runtime bytes.

**Subsystem affected**

The runner full-stack E2E workflow and its Daytona image build contract
are affected.

**Current behavior**

The content key hashes the full runner package. A development-only edit
changes the key even though the Docker build context excludes that edit.

**Proposed behavior**

The Dockerfile copies an explicit runtime build closure. The content key
hashes the same closure and continues to include every source, manifest,
lockfile, protocol, toolchain, and pinned image input that can affect
runtime bytes.

**Reason and benefit**

The workflow can reuse verified images for test-only changes. A runtime
change still creates a new immutable key and image.

**Breaking changes**

None. This changes only paid-test image cache identity and Docker build
inputs.

## What Changed

- Replace broad runner and eval package copies with explicit build
inputs.
- Advance the Daytona image content schema to version 5.
- Hash the matching explicit TypeScript, protocol, script, manifest,
lockfile, and Rust closure.
- Add contract coverage for runtime inputs and development-only
exclusions.

## Verification

- Focused Daytona image contract tests passed: 6 of 6.
- Exact-head ordinary CI [run
33913366909](https://github.com/paperclipai/paperclip/actions/runs/33913366909)
passed every job.
- The PR policy check passed on [run 33913366951, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/33913366951).
- The one-cell paid [run
33916670340](https://github.com/paperclipai/paperclip/actions/runs/33916670340)
passed end to end.
- Image job 101165705592 built the explicit 6.33 MB context from exact
source revision `4bcfb3faa7694aad4ceca2193230d9693af6c9e0`.
- The workflow published content key
`3a3a8a19d2362263e972bead4427048c82a7da61dc203cd5c83aa40b88d90524` at
immutable digest
`sha256:a5b6f7517bc020ec2bae8075210d1a3f867284f4733042114528e19150ffac0a`.
- Cosign verified the image and recorded transparency log entry
2715972694.
- The sole `core-compatibility.legacy-codex.daytona.message-marker` cell
passed in job 101168063383.
- Campaign aggregation, immutable S3 history publication, and GitHub
Pages publication all passed.
- Full local test, build, and typecheck suites were not run.

## Risks

A future Docker build input could be omitted from the explicit closure.
Contract tests reject the prior broad copies and check the current
required runtime inputs. The real Daytona image build also qualified the
closure before merge.

## Model Used

OpenAI Codex GPT-5.6 with agentic reasoning and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run focused tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 15:48:05 -05:00
Dotta 5716fe907e
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner subsystem executes agent work across local and managed
provider backends.
> - The lower pull requests restore the task runtime, provider backends,
and managed-provider control plane.
> - The restored system needs repeatable full-stack checks before it can
ship safely.
> - Paid live checks also need clear access, cost, and secret controls.
> - This pull request adds acceptance, live evaluation, chaos, and
release gates for the restored runner stack.
> - The benefit is measurable runner parity with safer release
decisions.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting. This change covers runner tests, release workflows,
server contracts, and evaluation tools.

**Problem or motivation**

The runner stack did not have one complete acceptance surface for native
Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could
miss provider drift, task-view regressions, cost-policy errors, and
destructive cleanup errors.

**Proposed solution**

Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid
workflows. Add live evaluation, chaos, cost-limit, redaction, and
release contract checks. Add AWS AgentCore infrastructure and guarded
provisioning tools. Keep the native runner experimental flag off by
default.

**Alternatives considered**

We considered manual smoke tests only. They do not give repeatable
evidence and they do not protect release branches. We also considered
one large pull request. The stacked pull requests keep each review below
the Greptile file limit.

**Roadmap alignment**

This work supports the shipped Cloud / Sandbox agents milestone and the
shipped Agent evals & feedback milestone in `ROADMAP.md`.

Related stack:

- #12699 adds managed provider backends and lifecycle support.
- #12691 adds qualified OpenCode and ACPX provider backends.
- #12685 restores task runtime rendering and steering.

## What Changed

- Add the runner full-stack harness with 57 catalog cells and 60 unit
tests.
- Add a Daytona runner image with digest-pinned base images and
base-aware image-content checks.
- Add guarded live evaluation and chaos workflows with a fixed
40-execution matrix; live and full-stack paid schedules now run only on
Sundays or by manual dispatch.
- Add in-flight reported-usage cost stops, post-turn cost caps,
exact-threshold failure classification, secret redaction, retry
classification, and actor authorization.
- Reattach stream and hard-budget listeners before restart-recovery
continuations so restored paid sessions cannot bypass in-flight
interruption.
- Preserve OpenCode usage and cost across tool-loop messages and turns
while exposing an explicit current-run delta to durable accounting.
- Keep PNG/WebM evidence in access-controlled artifacts only, reject
SVG, and publish only pruned inert structured per-attempt evidence.
- Add AWS AgentCore infrastructure, provisioning checks, and smoke
tools; reject unsafe model identifiers, require exact stack ownership
markers, and make failed-stack replacement explicit.
- Add evaluation-session contracts and capability reports.
- Add release workflow checks for immutable action pins, frozen
dependency installs, exact weekly cron shape, paid-run guards,
provider-secret isolation, and chaos test paths.
- Reauthorize the original and triggering numeric actor IDs as the first
step of every provider-secret job, including partial reruns, before
checkout or provider access.
- Give each full-stack matrix cell only its matching provider
credential, expose Daytona only to Daytona cells, and disable shared
dependency caches anywhere paid credentials or OIDC write access are
present.
- Protect the legacy manual E2E workflow with the same default-branch,
allowlist, environment, and per-job authorization boundary.
- Rotate live-eval candidates by week and retain 120 days of compatible
history so the seven-week trend window remains viable.
- Restore the root runner-acceptance commands and reconcile reported
snapshots,
raw receipts, and terminal usage without double counting or losing late
usage.
- Mark ACPX token deltas exact only when every budget field is present,
keep
cumulative cost/request authority separate, reject non-USD cost
labeling,
  and include thought tokens in output-token budgets.
- Keep `enableNativeRunner` off by default. The acceptance harness
enables it only in its isolated test instance.

## Verification

Passed locally:

- `pnpm --filter @paperclipai/paperclip-runner typecheck`
- `pnpm test:runner-acceptance:typecheck`
- `pnpm test:runner-acceptance` (19 tests)
- focused OpenCode proxy, driver, runnerd transport, live-session, and
turn-stream tests (106 tests)
- `pnpm --filter @paperclipai/paperclip-runner exec vitest run
src/live/clean-room-server.test.ts` (22 tests)
- `pnpm test:e2e:runner:typecheck`
- `pnpm test:e2e:runner:unit` (62 tests)
- `node --test scripts/__tests__/release-verify-workflow.test.mjs`
- `pnpm --filter @paperclipai/paperclip-runner
test:runner-workflow-evals` (22 tests)
- `pnpm -r typecheck`
- `pnpm build`
- `node --test
packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs`
(6 tests)
- `git diff --check`
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core
--lib --locked` (161 tests)
- focused ACPX provider-event tests (10 tests)
- The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged.

I did not run paid live provider jobs or provision AWS resources. Those
checks need credentials and can create cost.

## Risks

The paid workflows can create provider cost. They require an allowlisted
original and triggering actor, the protected `runner-e2e-paid`
environment, explicit opt-in variables, and cost limits. The four
provider credentials exist only in that master-only environment, which
requires allowlisted reviewer approval and disables administrator
bypass; repository and organization Actions scopes contain no copies.

Provider usage arrives after a billable request, so the live guard
cannot prevent one request from crossing a threshold. It interrupts
immediately on the first reported threshold hit and permits no
continuation.

Visual evidence can contain secrets rendered as pixels. PNG/WebM remain
only in access-controlled workflow artifacts; SVG and per-attempt XML
are excluded, and S3/Pages receive a pruned structured dashboard.

The AWS scripts can create cloud resources. They use explicit commands,
least-privilege roles, KMS encryption, saved nonsecret metadata, and
explicit teardown.

This pull request does not enable the experimental native runner for
existing instances.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5. The model used extended reasoning, tool use,
code execution, and parallel subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 08:55:08 -05:00