## Thinking Path
A fleet rollout failed with `image_manifest_not_found` for a commit that
had merged and gone green. Tracing that back: the fleet resolves
releases against the `-cloud` image, that image comes from `docker.yml`,
and `docker.yml`'s runs on master have been reading `cancelled` for a
long stretch. The cloud half was fine; the production half was hanging
and taking the run down with it — and, because a run holds the
concurrency slot for its whole duration, starving later commits of a
build at all.
## Linked Issues or Issue Description
No tracking issue — described inline, per CONTRIBUTING.md.
**What's wrong.** `build-and-push` builds `linux/amd64,linux/arm64` on
an x86 runner, so arm64 runs under QEMU. It hangs there —
deterministically, in the same step:
```
#111 [linux/arm64 build 8/10] RUN pnpm --filter @paperclipai/server build
```
…then emits nothing until `timeout-minutes: 60` kills it. Three
consecutive runs on 2026-09-04, silent for **38, 43 and 45 minutes**
respectively. The amd64 leg reached `production 5/5` minutes earlier in
every one.
**Why it stayed hidden.** A timed-out job is reported by GitHub as
**cancelled, not failed**. The run conclusion reads "cancelled", which
looks like supersession rather than breakage, so the production image
quietly stopped publishing.
**The knock-on.** A run that burns the full hour holds the top-level
concurrency slot for that hour. `cancel-in-progress: false` keeps
exactly one pending slot, so merges arriving faster than one an hour
supersede each other while queued. Sampling the last ten master commits,
**five produced no image at all** — their Docker runs have zero job
records because they never started.
**Expected.** Both architectures publish, and a commit merged during a
busy period still gets an image.
## What Changed
`build-and-push` becomes a two-leg matrix, each on a runner of its own
architecture:
| platform | runner |
|---|---|
| `linux/amd64` | `ubuntu-latest` |
| `linux/arm64` | `ubuntu-24.04-arm` |
Each leg pushes **by digest** (`push-by-digest=true`, untagged), and a
new `merge-and-push` job names the digests into one manifest list with
the real lane tags. Nothing is publicly tagged until the merge, so a
half-published multi-arch image is never a pullable state.
Two supporting changes:
- **Per-arch BuildKit cache refs** (`:buildcache-amd64` /
`:buildcache-arm64`). Separate runners sharing one ref would overwrite
each other on every build.
- **The PID-1 orphan-reaping check moves to the merge job**, since that
is where a tagged, pullable image first exists. It still runs against
the pushed image rather than a local build, for the same reason as
before.
**arm64 is kept, not dropped.** The cloud variant is amd64-only and can
be — managed hosts are amd64. This is the self-hosted image and ARM
hosts consume it, so dropping arm64 would break them. GitHub-hosted
arm64 runners are free for public repositories, which this is.
`build-and-push-cloud` is untouched. It was already `platforms:
linux/amd64` and has been succeeding in ~14 minutes throughout — that is
why `-cloud` images exist at all.
## Verification
Parsed the workflow and asserted its shape (jobs, matrix, `needs`, step
order, that the cloud job is unchanged). The artifact actions are pinned
by SHA with version comments, matching the repo's dominant convention —
`upload-artifact` v7 and `download-artifact` v8, the same pins used
across the other workflows; v8 is required for the `pattern` /
`merge-multiple` inputs the merge job uses.
**This PR's CI does not exercise the change.** `docker.yml` triggers on
master and tag pushes, never on pull requests — deliberately, since it
publishes release images. The first real run is after merge, so the
check is: the next master push produces a `Docker` run whose
`build-and-push (amd64)`, `build-and-push (arm64)` and `merge-and-push`
jobs all succeed, and whose conclusion is `success` rather than
`cancelled`.
## Risks
- **Not testable before merge**, per above. If the matrix is wrong the
next master push fails loudly rather than silently — which is already
better than the current state, where the failure mode is an invisible
"cancelled".
- **First use of `ubuntu-24.04-arm` in this repo.** No other workflow
uses an ARM runner. They are free for public repos, but if the label is
unavailable the arm64 leg will fail to schedule and the merge will not
run — no image, same as today, and visible.
- **Digest-push changes the publish shape.** Between the legs finishing
and the merge running, digests exist untagged in ghcr. Anything watching
for tags sees no intermediate state; anything enumerating untagged
manifests will see more of them.
- **Cache refs change name**, so the first build after this lands is
cold on both legs and will be slower than steady state.
- **Does not fix the underlying QEMU hang** — it avoids it. If arm64
ever has to build under emulation again, the same stall is presumably
still there.
- No application code, schema, server or persistence change.
## Model Used
Anthropic Claude — Opus 5, model ID `claude-opus-5`, run through Claude
Code.
Extended thinking enabled. Tool use throughout: GitHub Actions API to
correlate run/job outcomes and read build logs, `git` for ancestry
checks, and a YAML parser to validate the rewritten workflow's
structure.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
A workflow change has no unit test to add, and `docker.yml` cannot run
on a PR; the verification section states what to check on the first
master run instead.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Devin Foley <devin@paperclip.ing>