paperclip/packages/plugins/sandbox-providers/kubernetes
Nicky Leach 9edde68373
feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - AI agents run in sandboxed execution environments (Kubernetes pods,
Daytona workspaces, etc.) and need to sync files between the host and
those environments — for workspace setup, asset delivery, and output
retrieval
> - The existing sync path for Kubernetes uses a base64-over-exec chunk
loop: each ~4 MB chunk requires its own `execInPod` round-trip, so large
syncs balloon into many exec calls with corresponding overhead
> - `execInPod` supports piped stdin/stdout, meaning the full transfer
can be done as a single exec that streams a raw `tar` archive over the
data channel — one round-trip regardless of file size, with nothing
base64-encoded and nothing buffered whole in memory on either side
> - PR-1 (#10013, merged) added the
`onEnvironmentSyncIn`/`onEnvironmentSyncOut` opt-in hook API to the
sandbox provider interface and documented the protocol; PR-2 (#10028,
merged) implemented these hooks for the Daytona provider
> - This pull request implements the same two lifecycle hooks in the
Kubernetes sandbox provider, so workspace/asset file sync streams
through one `execInPod` per operation instead of the chunk loop
> - The benefit is significantly fewer exec round-trips for large syncs
and flat memory use on both host and pod, with security properties
preserved: atomic replace, secret-mode enforcement, path confinement,
TOCTOU-safe snapshot, and member-confinement on host-assembled archives
from sandbox-authored tar output

## Linked Issues or Issue Description

This is the third and final PR in a sequential series:
- Refs #10013 — PR-1: opt-in sync hook API + provider docs (merged)
- Refs #10028 — PR-2: native file-sync lifecycle hooks for Daytona
provider (merged)

**Feature:** Native single-exec file-sync lifecycle hooks for the
Kubernetes sandbox provider.

*Motivation:* The existing Kubernetes sync path encodes files as base64
and loops over `execInPod` one chunk at a time (~4 MB per exec). For
large workspaces or asset sets this is slow and resource-intensive. The
Kubernetes `execInPod` API supports piped stdin/stdout, enabling a
raw-`tar` streaming transfer that needs only one exec regardless of file
count or size and never buffers the whole payload in memory.

*Proposed solution:* Implement `onEnvironmentSyncIn` and
`onEnvironmentSyncOut` in the Kubernetes provider using a streaming
`execInPod` with a tar pipeline — for syncIn the host builds the archive
on disk and streams its raw bytes into the pod's stdin (`head -c
<exact-size> | tar -x`, no base64); for syncOut in-pod `tar` writes to
the exec's stdout and the host streams those bytes straight to a file.
Path confinement, atomic replace, secret-mode enforcement, TOCTOU
protection, and a streamed-bytes fail-closed guard are all enforced.

## What Changed

- **New `src/file-sync.ts`** in
`packages/plugins/sandbox-providers/kubernetes/` — `performSyncIn` and
`performSyncOut` over an injected pod-exec closure, keeping transfer
logic hermetically unit-testable
- **New `execInPodStreaming` in `src/pod-exec.ts`** — a streaming exec
primitive that binds a caller-supplied stdin readable and a stdout
writable to the exec WebSocket data channel, added alongside the
existing `execInPod` (which is unchanged). This lets a transfer stream
raw bytes to/from disk instead of buffering the payload as a single
string
- **Updated `src/plugin.ts`** — registers
`onEnvironmentSyncIn`/`onEnvironmentSyncOut`; resolves the `sandbox-cr`
pod exactly like `onEnvironmentExecute` and delegates; `job` backend
rejects file-sync calls explicitly (out of scope)
- **syncIn path:** host builds the tarball to a temp file → streams its
raw bytes over exec stdin, bounded in-pod by `head -c
<exact-archive-size> | tar -x` (no base64 anywhere) → extract into a
`/proc/self/fd`-pinned reserved `0700` staging dir → `chmod`-before-`mv
-f` atomic replace per file (directory mappings use
`followSymlinks`→`-h`)
- **syncOut path:** in-pod validate + realpath-snapshot each source
(closes the validation→copy TOCTOU window) → single-exec `tar -c`
streamed over exec stdout → host streams that stdout straight to a temp
file through a byte-counting transform → member-confined extraction of
the sandbox-authored archive
- **Security properties:** secret files land at requested mode with no
widened window; every interpolated path is shell-quoted and confined
lexically plus via in-pod `realpath`; the outbound stream is bounded by
a **streamed-bytes disk guard** (`MAX_SYNC_OUTPUT_BYTES`, 8 GiB default,
per-call overridable) that fails the transfer closed — writing no target
file — if an untrusted pod emits more bytes than allowed. Neither host
nor pod buffers the whole payload, so there is no in-memory size cap on
the transfer
- **No changes** to `execInPod`, `wrapCommandWithEnv`, or
`FastUploadInterceptor` (the `environmentExecute` path is untouched)
- **No dependency or lockfile changes**
- **New tests** in `test/unit/file-sync.test.ts` (atomic-replace, `0600`
secret mode, symlink preserve/deref, dir-mapping, exclude,
path-confinement rejection, streamed-output guard fail-closed) and
`test/unit/pod-exec.test.ts` (streaming stdin/stdout, caller-sink error
fail-closed), plus extended `test/unit/plugin.test.ts`

## Follow-up: Legacy Job-Lease Base64 Fallback Fix

Addresses the Greptile 4/5 blocking finding ("Handle existing job
leases", `server/src/services/environment-runtime.ts`).

Job leases provisioned before the `nativeFileSyncUnsupported` metadata
flag existed carry `backend: "job"` but no flag, so `supportsSync()`
treated them as native-capable and routed their sync to the pod-exec
hook — which the job backend rejects (it has no exec channel) instead of
using the byte-identical base64 fallback. The fix adds a
belt-and-suspenders gate on the persisted `backend === "job"` field
alongside the existing `nativeFileSyncUnsupported` flag check, so
pre-existing job leases continue syncing via the base64 fallback after
deployment. No behaviour change for `sandbox-cr` leases.

## Verification

- `pnpm --filter @paperclipai/sandbox-provider-kubernetes test` — 19
files / 182 tests green, including the existing `upload-interceptor` and
`pod-exec` suites
- `tsc --noEmit` in the kubernetes package — 0 errors
- The sync hooks are opt-in; existing `environmentExecute` behaviour is
unaffected and tested by the unchanged existing suites

## Risks

- **Opt-in only:** `onEnvironmentSyncIn`/`onEnvironmentSyncOut` are
registered conditionally; providers that do not register them fall back
to the existing chunk loop. No regression risk on the existing path.
- **Shell-injection surface:** all path interpolation uses
shell-quoting; paths are additionally confined lexically and via in-pod
`realpath` before use.
- **TOCTOU on syncOut:** the in-pod snapshot validates and records file
metadata before the tar call, closing the window between validation and
copy.
- **Archive member confinement:** host-side reassembly rejects any tar
member whose resolved path escapes the target directory, preventing a
malicious in-pod tar from writing outside the intended destination.
- **Untrusted-output volume:** an over-large outbound stream trips the
streamed-bytes disk guard and fails closed (no target written and the
temp sink is swept) rather than filling host disk or memory; the guard
bounds disk unconditionally and bounds memory insofar as WebSocket
write-backpressure holds.

## Model Used

Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`) — produced by a
Claude-based AI agent using agentic tool use and multi-step code
generation. 200K context window, extended reasoning, code execution and
verification capabilities.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Harold Kim <harold@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-22 18:52:47 -07:00
..
manifests
…
src feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053) 2026-07-22 18:52:47 -07:00
test feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053) 2026-07-22 18:52:47 -07:00
.gitignore
…
README.md Make ACP the default engine for local adapters (#9238) 2026-07-08 19:05:03 -07:00
SMOKE.md
…
package.json fix(plugins): move dev SDK linking out of plugin postinstall scripts (#8255) 2026-06-18 07:45:53 -07:00
tsconfig.json
…
vitest.config.ts
…

README.md

@paperclipai/plugin-kubernetes (alpha)

First-party Paperclip sandbox-provider plugin for Kubernetes.

Alpha: the default backend (sandbox-cr) is built on kubernetes-sigs/agent-sandbox v1alpha1 — expect breaking changes as that CRD evolves toward Beta. A stable fallback backend (job, using batch/v1 Job) is available for clusters without agent-sandbox installed, but it does NOT support multi-command exec (paperclip-server's adapter-install pattern requires sandbox-cr).

Prerequisites

  1. A Kubernetes cluster running k8s 1.27+
  2. kubernetes-sigs/agent-sandbox controller installed in the cluster (alpha — installs the sandboxes.agents.x-k8s.io/v1alpha1 CRD and controller)
  3. Paperclip-server running with access to the cluster (in-cluster via inCluster: true or external via kubeconfig)

For job backend (stable fallback)

  1. A Kubernetes cluster running k8s 1.27+
  2. Paperclip-server with cluster access — no additional controllers or CRDs required

Installation

paperclipai plugin install @paperclipai/plugin-kubernetes

Or, for local development:

paperclipai plugin install --local /path/to/paperclip/packages/plugins/sandbox-providers/kubernetes

Backends

The plugin supports two backend modes, selected via the backend config field:

Backend Default Stability Multi-command exec Requires
sandbox-cr Yes Alpha Yes kubernetes-sigs/agent-sandbox controller
job No Stable No Nothing beyond k8s 1.27+

sandbox-cr (default): Creates a Sandbox CR (agents.x-k8s.io/v1alpha1) whose controller provisions a long-lived pod running sleep infinity. paperclip-server execs individual commands into the running pod — this is the multi-command adapter-install pattern. When you releaseLease, the Sandbox CR is deleted and the controller tears down the pod.

job (stable fallback): Creates a batch/v1 Job. The container entrypoint runs once and exits — no multi-command exec possible. Use this when you cannot install agent-sandbox, or when you need strictly stable Kubernetes APIs. Note: paperclip-server's adapter-install pattern will not work in job mode.

Migrating from job to sandbox-cr

  1. Install the agent-sandbox controller: kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/latest/download/install.yaml
  2. Update your environment config to set backend: "sandbox-cr" (or remove backend since sandbox-cr is the default)
  3. New leases will use the Sandbox CR backend. Existing leases created with job mode continue to use job semantics until they are released.

Configuration

Create a sandbox environment with driver: kubernetes. One of these auth fields is required:

  • inCluster: true — use the in-pod ServiceAccount credentials (when paperclip-server runs inside the same cluster).
  • kubeconfig: <YAML> — inline kubeconfig (stored as a company secret).
  • kubeconfigSecretRef: <secret-uuid> — reference to an existing Paperclip secret.

Common optional fields:

Field Default Purpose
backend "sandbox-cr" sandbox-cr (alpha, requires agent-sandbox controller) or job (stable, one-shot entrypoint).
adapterType "claude_local" One of the supported adapter types (claude_local, codex_local, gemini_local, cursor_local, opencode_local, pi_local). Determines runtime image + env keys + egress allow-list.
namespacePrefix "paperclip-" Prefix for the per-company tenant namespace.
companySlug derived from companyId Override the auto-derived company slug.
imageRegistry (none) Override the default registry for agent runtime images.
imageAllowList [] Glob patterns of allowed target.imageOverride values. Empty = no override permitted.
imagePullSecrets [] Names of pre-created Docker image pull secrets in the tenant namespace.
egressAllowFqdns [] Additional FQDNs (beyond adapter defaults like api.anthropic.com).
egressAllowCidrs [] Additional CIDRs to allow egress to.
egressMode "standard" standard (NetworkPolicy + CIDRs) or cilium (CiliumNetworkPolicy + FQDN allow-list).
runtimeClassName (none) e.g. kata-fc for Firecracker-backed microVMs. Cluster must have the RuntimeClass installed.
serviceAccountAnnotations {} Annotations applied to per-tenant ServiceAccount (e.g. IRSA eks.amazonaws.com/role-arn).
jobTtlSecondsAfterFinished 900 Seconds after a Job completes before garbage-collection.
podActivityDeadlineSec 3600 Hard ceiling on a single run's wall-clock time.

Full JSON Schema in src/manifest.ts.

What gets created in your cluster

For each company that runs agents (created lazily on first dispatch):

Namespace          paperclip-{companySlug}        (PSS: restricted enforce + audit)
ServiceAccount     paperclip-tenant-sa
Role               paperclip-tenant-role          (only get pods/log)
RoleBinding        paperclip-tenant-rb
ResourceQuota      paperclip-quota                (pods, requests/limits cpu+memory)
LimitRange         paperclip-limits               (container max/min/default/defaultRequest)
NetworkPolicy      paperclip-deny-all             (deny ingress + egress baseline)
NetworkPolicy      paperclip-egress-allow         (DNS + paperclip-server callback + user CIDRs)
                   OR CiliumNetworkPolicy paperclip-egress-fqdn if egressMode=cilium

For each agent run (sandbox-cr backend):

Sandbox CR         pc-{ulid}                       (agents.x-k8s.io/v1alpha1; explicit delete on release)
Pod                pc-{ulid}-{podSuffix}           (managed by Sandbox controller; torn down on CR delete)
Secret             pc-{ulid}-env                   (owned by Sandbox CR; cascade-deleted)

For each agent run (job backend):

Job                pc-{ulid}                       (backoffLimit: 0, ttlSecondsAfterFinished from config)
Pod                pc-{ulid}-{podSuffix}           (owned by Job; cascade-deleted)
Secret             pc-{ulid}-env                   (owned by Job; cascade-deleted)

Security baseline

Every agent pod is:

  • non-root (runAsUser: 1000, runAsGroup: 1000, runAsNonRoot: true)
  • drops ALL Linux capabilities, allowPrivilegeEscalation: false
  • readOnlyRootFilesystem: true with explicit emptyDir mounts for /workspace, /home/paperclip, /home/paperclip/.cache, /tmp
  • seccompProfile: RuntimeDefault
  • Tini as PID 1 (reaps zombies, forwards signals)
  • fsGroupChangePolicy: OnRootMismatch (fast PVC startup; openclaw-operator lesson)
  • automountServiceAccountToken: true (for the agent shim's paperclip-server callback)

Plus per-namespace pod-security.kubernetes.io/enforce: restricted and a deny-all NetworkPolicy baseline with explicit egress allow-list (DNS, paperclip-server, configured FQDNs/CIDRs).

The per-run Secret carrying the bootstrap token and adapter API keys has ownerReferences pointing at the owning Job, so a single kubectl delete job … cascades cleanly to the Pod and Secret.

Optional Kata-FC microVM isolation

For stronger isolation, install Kata Containers with the Firecracker hypervisor, then set runtimeClassName: kata-fc in the plugin config. Each agent pod will run inside a Firecracker microVM. Requires nested-virt-capable nodes (bare-metal or specific cloud instance types).

Roadmap

  • Phase A (done): sandbox-cr backend — multi-command exec via agent-sandbox Sandbox CRD.
  • Phase B: Warm pool support — pre-provisioned Sandbox CRs for sub-second cold starts. The SandboxOrchestrator interface reserves optional pause?/resume? extension slots.
  • Phase C: Kata-FC + snapshots — runtimeClassName: kata-fc with VM snapshot for fast restore.
  • Phase D: Contribute back to agent-sandbox upstream if their Beta model diverges from our needs. The SandboxOrchestrator interface (src/sandbox-orchestrator.ts) is the clean swap point — a new implementation can be added without touching plugin.ts business logic.

Lessons learned (from openclaw-operator)

This plugin adopts patterns from openclaw-rocks/openclaw-operator:

  • Tini PID 1 (issue #471 — zombie helper processes)
  • Read-only rootFS with explicit writable mounts (issue #456 — ~/.config not writable)
  • Strategic merge on reconcile (issue #446 — preserve third-party annotations)
  • Multi-storage-class testing (issue #448 — local-path-provisioner differences)
  • Image version compat matrix (issue #462 — runtime deps cannot resolve after upgrade)

Development

cd packages/plugins/sandbox-providers/kubernetes
pnpm install --ignore-workspace
pnpm test           # unit tests only (fast)
pnpm typecheck
pnpm build

To run the kind-cluster integration test (requires kubectl --context kind-paperclip and a pre-loaded alpine image; see test/integration/end-to-end-run.test.ts):

RUN_K8S_INTEGRATION_TESTS=1 pnpm test test/integration/end-to-end-run.test.ts