## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-backed runs need observable startup behavior so operators can see where time is spent before an adapter is invoked > - The current startup path only surfaced aggregate timing, which makes it hard to identify the slow boundary in the bring-up sequence > - That gap matters because sandbox startup latency is often dominated by one specific step, and aggregate timing hides the bottleneck > - This pull request adds per-step startup timing events for the named sandbox bring-up boundaries > - The benefit is more precise observability with no control-flow change and no schema migration ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (multiple of the above) ### Problem or motivation Sandbox run startup only exposed aggregate timing. That makes it hard to identify which bring-up boundary is responsible for slow starts, especially in remote or sandboxed execution where the bottleneck can move between workspace setup, skill reconciliation, bridge setup, and adapter handshake. ### Proposed solution Emit a structured timing event for each named startup boundary before the adapter is invoked, so the existing run-event stream carries per-step duration data. This keeps the event path additive and lets operators see which step dominates startup latency without changing control flow or introducing a schema migration. ### Alternatives considered - Keep only the aggregate startup duration: simpler, but it hides the bottleneck and makes regression analysis much harder. - Add a new telemetry sink or schema field: rejected because the existing run-event payload already carries structured event data and does not need a new storage path. - Log unstructured text for each step: rejected because it is harder to query and aggregate than a structured `step` + `durationMs` event. ### Roadmap alignment This fits the roadmap direction around cloud / sandbox agents and enforced outcomes by improving observability for sandboxed execution without changing the control plane model. The roadmap section is broad, but it does not call out this specific startup-timing work as a planned duplicate. ### Additional context This PR is intentionally additive. It records timing for the named startup boundaries in the existing event stream and leaves the bridge, database shape, and adapter invocation order unchanged. ## What Changed - Added a `measureStartupStep` helper that times a startup step, emits one structured `run.startup.step` event, and rethrows failures after recording duration - Wrapped the seven sandbox bring-up boundaries in `execute.ts` so the structured timing covers each named step before adapter invocation - Added unit coverage for the helper and integration coverage for the startup-step events in the adapter-utils execute path - Kept the event path additive, with no bridge change and no database migration ## Verification - `tsc --noEmit` for `@paperclip/adapter-utils` - `pnpm test` in `packages/adapter-utils` equivalent suite coverage: 292 passed, 4 skipped - Adjacent server event/log-store suites: `run-log-store.test.ts` and `heartbeat-run-log.test.ts` passed (11 total) - Git validation: fetched `origin/feat/sandbox-startup-step-timing`, confirmed it matches the authorized submit SHA, and confirmed `origin/master..origin/feat/sandbox-startup-step-timing` contains the expected single commit - Searched GitHub for duplicate or related open PRs/issues and found no overlapping open items - Checked `ROADMAP.md`; the roadmap covers sandboxed environments generally, but does not call out this specific startup-timing observability work as a planned duplicate ## Risks - Low risk: the change is additive and only emits additional structured events - If downstream consumers assume startup events are aggregate-only, they may need to ignore or account for the new `run.startup.step` entries - Timing is measured via the injected clock and event emission happens in a `finally`, so failures still report duration before rethrowing ## Model Used OpenAI GPT-5, tool-using coding agent ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|---|---|---|
| .. | ||
| src | ||
| CHANGELOG.md | ||
| README.md | ||
| package.json | ||
| tsconfig.json | ||
README.md
@paperclipai/adapter-utils
Shared utilities for Paperclip adapters: process spawning, environment injection, sandbox/SSH transport, workspace sync, and the round-trip helpers that move code between the local execution-workspace cwd and wherever the agent actually runs.
For the adapter-author guide see
docs/adapters/creating-an-adapter.md
and the in-repo notes at packages/adapters/AUTHORING.md.
No-remote-git contract
The local execution-workspace cwd is the only persistence boundary across runs. No adapter may depend on a git remote for cross-run state.
Adapters that run the agent on a different host should use the SSH round-trip
helpers in src/ssh.ts:
prepareWorkspaceForSshExecution({ spec, localDir, remoteDir })— bundles the local cwd (tracked files, dirty edits, untracked additions, and the git history needed to reconstruct it) toremoteDirbefore the run starts. Runs with nogit remoteconfigured.restoreWorkspaceFromSshExecution({ spec, localDir, remoteDir, ... })— syncs the remote cwd back intolocalDirafter the run, including any new commits the agent created. Also runs with nogit remoteconfigured.
prepareRemoteManagedRuntime in
src/remote-managed-runtime.ts wraps both
calls for adapters that want a per-run remote workspace and an automatic
restoreWorkspace() finally hook.
The invariant is pinned by the no-remote-git contract case in
src/ssh-fixture.test.ts, which asserts that a
remote-only commit propagates to the local worktree through the
prepare → restore round-trip with no git remote configured at any point. Do
not regress that test.