fix(server,ui,claude-local): verify auth on the adapter Test lane and enforce managed-sandbox tenant binding (#11810)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter Test checks whether an agent adapter can run with its configured environment, and every local-driver adapter (Claude, Codex, Gemini, OpenCode, Pi, Cursor, etc.) shares this Test route and its UI resolution logic > - The Claude ACP Test lane could report pass without checking local or remote authentication, and the shared Test route and UI had gaps in environment binding, probe safety, and managed-sandbox resolution that affect every adapter that uses the Test button, not only Claude > - This pull request verifies authentication on every Claude ACP target, and closes the shared Test-route/UI gaps: tenant-binding on the route, a managed-sandbox-only redirect that matches the real run path, and a three-tier environment resolution in the UI > - The benefit is a truthful Test result with safer probe execution and tenant isolation, for Claude specifically and for every other local adapter that shares this Test surface ## Linked Issues or Issue Description **What happened?** The Claude ACP Test lane returned `status: "pass"` without checking authentication for some local and non-sandbox targets. Separately, the shared `/companies/:companyId/adapters/:type/test-environment` route — used by every local-driver adapter, not only Claude — accepted a foreign environment id, and its UI resolution did not mirror the server's managed-sandbox-only redirect. **Expected behavior** The Test lane checks the resolved credential and hello probe for every Claude ACP target. The shared adapter Test route rejects a foreign environment before it reveals environment details or starts a lease, for any adapter type. The Test's environment resolution (UI and server) matches the real run's three-tier resolution, including the managed-sandbox-only redirect. **Steps to reproduce** 1. Run the Claude ACP Test lane against a local target without a valid credential. 2. Run the adapter Test route with an environment id from another company (any adapter type). 3. Observe the pass result on step 1, or the missing tenant-binding rejection on step 2. **Paperclip version or commit** `933749e01f74e82ce5d315c071be534d04e01158` **Deployment mode** Local dev (`pnpm dev`) and server route tests. **Agent adapter(s) involved** Claude Code directly (the ACP auth-verification work). The tenant-binding guard, managed-sandbox-only redirect, and UI three-tier resolution apply to the shared adapter Test route and affect every local-driver adapter (Codex, Gemini, OpenCode, Pi, Cursor, etc.), not only Claude — see "What Changed" below for the split between Claude-only and shared changes. **Database mode** Not database-related. **Access context** Both board and agent paths use the affected Test surface, for every local-driver adapter. **Additional context** Two commits that were previously bundled into this PR — a `plugin-worker-manager` duplex-channel frame-bound fix and a `workspace-runtime` exit-persist crash fix — are unrelated to the adapter Test lane and have been split out into their own PRs: #11860 and #11861. ## What Changed Claude-only (`packages/adapters/claude-local`): - Verify `CLAUDE_CODE_OAUTH_TOKEN` and run the hello probe for every Claude ACP target. - Keep `adapter_auth_missing` sandbox-only and report missing non-sandbox credentials as a warning. - Add a deny-by-default probe environment builder for the ACP and CLI local probes. - Log only fixed probe context and allowlisted classifications. - Seed the host OAuth token into the hello probe environment. Shared, cross-adapter (`server/src/routes/agents.ts`, `ui/src/lib/adapter-test-environment.ts`, `ui/src/components/AgentConfigForm.tsx`, `ui/src/components/OnboardingWizard.tsx`): - Add a company-binding guard and a binding assertion for the generic `/companies/:companyId/adapters/:type/test-environment` route, so a foreign-company environment id is rejected before any secret resolution or sandbox lease, for every adapter type. - Resolve all three server environment tiers (agent default, instance default, local default) in the UI, and add the managed-sandbox-only redirect so the Test probes the same target a real run would use. - Enforce onboarding Test results: block hire on a failed environment test. - Add regression tests for authentication, tenant binding, probe safety, diagnostics, and UI resolution. ## Verification - Adapter suites pass for the Claude local server probe, remote, ACP, auth, probe environment, and config paths. - Server route tests pass, including the five tenant-binding cases. - UI adapter Test environment resolver tests pass for all three resolution tiers. - Adapter package `tsc --noEmit` exits 0. - Full CI must pass on this pull request. ## Risks The probe environment now denies caller variables by default. A required variable that is not on the allowlist could stop a probe from starting. The route now rejects foreign environment ids with a fixed 403 response. The managed-sandbox-only redirect changes where the Test (and the login affordance) probes for every local-driver adapter under that policy, not only Claude — operators running other local adapters under managed-sandbox-only will see their Test target move from local to the managed sandbox, matching what real runs already do. The change limits secret and diagnostic exposure. ## Model Used Original implementation: OpenAI Codex, GPT-5; exact context window not exposed in that run; tool use and code execution. This revision (commit split and title/description correction): Claude, Sonnet 5 (claude-sonnet-5). The original title and description described this PR as Claude-only; review found it also changes the shared adapter Test route and UI resolution used by every local-driver adapter, and carried two unrelated server fixes. Claude split those two commits into #11860 and #11861 via `git rebase --onto` (verified byte-identical to the original tree minus those commits) and rewrote this description to reflect the actual scope. No functional code in this PR was authored by Claude. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
parent
917d2350f8
commit
5bc6031f79
|
|
@ -1,4 +1,7 @@
|
|||
import { afterEach, describe, expect, it, vi } from "vitest";
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { chmod, mkdtemp, rm, writeFile } from "node:fs/promises";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import type { AdapterExecutionResult } from "@paperclipai/adapter-utils";
|
||||
import type { AdapterExecutionTarget } from "@paperclipai/adapter-utils/execution-target";
|
||||
|
||||
|
|
@ -39,7 +42,11 @@ vi.mock("@paperclipai/adapter-utils/execution-target", async () => {
|
|||
};
|
||||
});
|
||||
|
||||
import { mapClaudeAcpAuthErrorCode, probeClaudeAcpSandboxLogin } from "./acp.js";
|
||||
import {
|
||||
mapClaudeAcpAuthErrorCode,
|
||||
probeClaudeAcpSandboxLogin,
|
||||
testClaudeAcpEnvironment,
|
||||
} from "./acp.js";
|
||||
import { ADAPTER_AUTH_MISSING_CHECK_CODE } from "./auth-check.js";
|
||||
|
||||
const sandboxTarget: AdapterExecutionTarget = {
|
||||
|
|
@ -310,11 +317,15 @@ describe("probeClaudeAcpSandboxLogin", () => {
|
|||
expect(checks[0]?.level).toBe("warn");
|
||||
});
|
||||
|
||||
it("never copies a thrown probe error into a Test-result check", async () => {
|
||||
// A sandbox transport failure can carry a credential. Inject a secret
|
||||
// marker through the thrown error and assert no check text repeats it.
|
||||
const secret = "sk-ant-LEAKMARKER0123456789abcdef";
|
||||
probeResult.throwError = new Error(`transport failed with ${secret}`);
|
||||
it("never copies a thrown probe error into a Test-result check or the log", async () => {
|
||||
// A sandbox transport failure can carry a credential. Inject an opaque
|
||||
// credential marker and a proxy marker through the thrown error, then assert
|
||||
// no check text and no log call repeats either one.
|
||||
const opaqueCredMarker = "OPAQUECREDMARKERnoshape";
|
||||
const proxyMarker = "http://user:pass@proxy.corp.internal:3128";
|
||||
probeResult.throwError = new Error(
|
||||
`transport failed with ${opaqueCredMarker} via ${proxyMarker}`,
|
||||
);
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
|
||||
const checks = await probeClaudeAcpSandboxLogin({
|
||||
|
|
@ -323,25 +334,32 @@ describe("probeClaudeAcpSandboxLogin", () => {
|
|||
});
|
||||
|
||||
const checkText = JSON.stringify(checks);
|
||||
expect(checkText).not.toContain(secret);
|
||||
expect(checkText).not.toContain("LEAKMARKER");
|
||||
expect(checkText).not.toContain(opaqueCredMarker);
|
||||
expect(checkText).not.toContain("proxy.corp.internal");
|
||||
expect(checks[0]?.code).toBe("claude_acp_login_probe_unavailable");
|
||||
// The diagnostic still reaches the server log, but the secret is redacted.
|
||||
// The log carries only the fixed context, the allowlisted classification,
|
||||
// and a safe error class name. It never repeats the raw error text.
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(secret);
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
expect(loggedText).not.toContain(opaqueCredMarker);
|
||||
expect(loggedText).not.toContain("proxy.corp.internal");
|
||||
expect(warnSpy.mock.calls[0]?.[1]).toMatchObject({
|
||||
classification: "spawn_error",
|
||||
errorClass: "Error",
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("never copies raw probe stderr or stdout into a Test-result check", async () => {
|
||||
// A non-zero probe can print a credential to stderr. Inject a secret marker
|
||||
// and assert no check text repeats it.
|
||||
const secret = "sk-ant-STDERRMARKER0123456789abcdef";
|
||||
it("never copies raw probe stderr or stdout into a Test-result check or the log", async () => {
|
||||
// A non-zero probe can print an opaque credential to stdout and a proxy URL
|
||||
// to stderr. Inject one marker in each stream and assert neither reaches a
|
||||
// check or the log.
|
||||
const opaqueCredMarker = "OPAQUECREDMARKERstream";
|
||||
const proxyMarker = "http://user:pass@proxy.corp.internal:3128";
|
||||
probeResult.value = {
|
||||
exitCode: 2,
|
||||
stdout: initLine,
|
||||
stderr: `fatal: leaked ${secret}`,
|
||||
stdout: [initLine, `note: ${opaqueCredMarker}`].join("\n"),
|
||||
stderr: `fatal: proxy connect failed ${proxyMarker}`,
|
||||
timedOut: false,
|
||||
};
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
|
|
@ -352,14 +370,276 @@ describe("probeClaudeAcpSandboxLogin", () => {
|
|||
});
|
||||
|
||||
const checkText = JSON.stringify(checks);
|
||||
expect(checkText).not.toContain(secret);
|
||||
expect(checkText).not.toContain("STDERRMARKER");
|
||||
expect(checkText).not.toContain(opaqueCredMarker);
|
||||
expect(checkText).not.toContain("proxy.corp.internal");
|
||||
expect(checks[0]?.code).toBe("claude_acp_login_probe_unavailable");
|
||||
// The diagnostic still reaches the server log, but the secret is redacted.
|
||||
// The log carries only the fixed context, the allowlisted classification,
|
||||
// and the safe exit code. It never repeats the raw stream text.
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(secret);
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
expect(loggedText).not.toContain(opaqueCredMarker);
|
||||
expect(loggedText).not.toContain("proxy.corp.internal");
|
||||
expect(warnSpy.mock.calls[0]?.[1]).toMatchObject({
|
||||
classification: "nonzero_exit",
|
||||
exitCode: 2,
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
});
|
||||
|
||||
describe("Claude ACP hello probe on local and SSH targets", () => {
|
||||
const sshTarget: AdapterExecutionTarget = {
|
||||
kind: "remote",
|
||||
transport: "ssh",
|
||||
remoteCwd: "/home/user/paperclip-workspace",
|
||||
spec: { host: "example.com", port: 22, username: "user" },
|
||||
} as unknown as AdapterExecutionTarget;
|
||||
|
||||
// Clear the host proxy and host auth variables so a local probe reads a
|
||||
// deterministic env regardless of the machine that runs the suite.
|
||||
const CLEARED_HOST_ENV_KEYS = [
|
||||
"HTTP_PROXY",
|
||||
"HTTPS_PROXY",
|
||||
"NO_PROXY",
|
||||
"http_proxy",
|
||||
"https_proxy",
|
||||
"no_proxy",
|
||||
"ANTHROPIC_API_KEY",
|
||||
"CLAUDE_CODE_OAUTH_TOKEN",
|
||||
"ANTHROPIC_AUTH_TOKEN",
|
||||
"CLAUDE_CONFIG_DIR",
|
||||
"CLAUDE_CODE_USE_BEDROCK",
|
||||
"ANTHROPIC_BEDROCK_BASE_URL",
|
||||
];
|
||||
let tempDir: string | null = null;
|
||||
let claudePath = "";
|
||||
let savedPath: string | undefined;
|
||||
let savedEnv: Record<string, string | undefined> = {};
|
||||
|
||||
beforeEach(async () => {
|
||||
tempDir = await mkdtemp(path.join(os.tmpdir(), "paperclip-acp-localprobe-"));
|
||||
claudePath = path.join(tempDir, "claude");
|
||||
await writeFile(claudePath, "#!/bin/sh\nexit 0\n");
|
||||
await chmod(claudePath, 0o755);
|
||||
savedPath = process.env.PATH;
|
||||
process.env.PATH = tempDir;
|
||||
savedEnv = {};
|
||||
for (const key of CLEARED_HOST_ENV_KEYS) {
|
||||
savedEnv[key] = process.env[key];
|
||||
delete process.env[key];
|
||||
}
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
process.env.PATH = savedPath;
|
||||
for (const [key, value] of Object.entries(savedEnv)) {
|
||||
if (value === undefined) delete process.env[key];
|
||||
else process.env[key] = value;
|
||||
}
|
||||
if (tempDir) await rm(tempDir, { recursive: true, force: true }).catch(() => undefined);
|
||||
tempDir = null;
|
||||
});
|
||||
|
||||
it("emits an explicit OAuth-token check on the ACP lane when CLAUDE_CODE_OAUTH_TOKEN is set", async () => {
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp", env: { CLAUDE_CODE_OAUTH_TOKEN: "oauth-token-secret" } },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
expect(result.checks.some((check) => check.code === "claude_oauth_token_configured")).toBe(true);
|
||||
// Every result names the target it probed, including the host case.
|
||||
expect(result.checks.some((check) => check.code === "claude_environment_target")).toBe(true);
|
||||
// The token value never enters a check.
|
||||
expect(JSON.stringify(result.checks)).not.toContain("oauth-token-secret");
|
||||
});
|
||||
|
||||
it("runs on a local target and reports auth-required without the sandbox-only adapter_auth_missing", async () => {
|
||||
probeResult.value = { exitCode: 1, stdout: loginRequiredStdout, stderr: "", timedOut: false };
|
||||
const checks = await probeClaudeAcpSandboxLogin({ config: { engine: "acp" }, target: null });
|
||||
expect(checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(true);
|
||||
expect(checks.some((check) => check.code === ADAPTER_AUTH_MISSING_CHECK_CODE)).toBe(false);
|
||||
});
|
||||
|
||||
it("runs on an SSH target and reports auth-required without adapter_auth_missing", async () => {
|
||||
probeResult.value = { exitCode: 1, stdout: loginRequiredStdout, stderr: "", timedOut: false };
|
||||
const checks = await probeClaudeAcpSandboxLogin({ config: { engine: "acp" }, target: sshTarget });
|
||||
expect(checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(true);
|
||||
expect(checks.some((check) => check.code === ADAPTER_AUTH_MISSING_CHECK_CODE)).toBe(false);
|
||||
});
|
||||
|
||||
it("spawns the trusted resolved claude and drops hostile caller env on a local probe", async () => {
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
process.env.HTTPS_PROXY = "http://trusted-proxy:8443";
|
||||
|
||||
await probeClaudeAcpSandboxLogin({
|
||||
config: {
|
||||
engine: "acp",
|
||||
env: {
|
||||
ANTHROPIC_API_KEY: "keep-this-key",
|
||||
NODE_OPTIONS: "--require /hostile/evil.js",
|
||||
PATH: "/hostile/bin",
|
||||
LD_PRELOAD: "/hostile/evil.so",
|
||||
HTTP_PROXY: "http://caller-proxy:8080",
|
||||
},
|
||||
},
|
||||
target: null,
|
||||
});
|
||||
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedCommand = call[2] as string;
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
// The trusted resolved claude executable, never the caller command path.
|
||||
expect(spawnedCommand).toBe(claudePath);
|
||||
// The approved key reaches the child; the hostile keys never do.
|
||||
expect(spawnedEnv.ANTHROPIC_API_KEY).toBe("keep-this-key");
|
||||
expect(spawnedEnv.NODE_OPTIONS).toBeUndefined();
|
||||
expect(spawnedEnv.PATH).toBeUndefined();
|
||||
expect(spawnedEnv.LD_PRELOAD).toBeUndefined();
|
||||
expect(spawnedEnv.HTTP_PROXY).toBeUndefined();
|
||||
// The trusted proxy reaches the child; the caller proxy never does.
|
||||
expect(spawnedEnv.HTTPS_PROXY).toBe("http://trusted-proxy:8443");
|
||||
expect(JSON.stringify(spawnedEnv)).not.toContain("caller-proxy");
|
||||
});
|
||||
|
||||
it("runs the host login probe with the host ANTHROPIC_API_KEY on a local target", async () => {
|
||||
// A local ACP run inherits the host environment, so a host ANTHROPIC_API_KEY
|
||||
// authenticates the real run. The Test lane runs the login probe with the
|
||||
// same host key, so the probe env matches the credential the real run
|
||||
// receives. The probe then reports a real result, not a false auth-required.
|
||||
process.env.ANTHROPIC_API_KEY = "sk-ant-host-key";
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp" },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
// The probe runs once with the host key, so it authenticates and reports no
|
||||
// false auth-required and no probe-unavailable check.
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
expect(spawnedEnv.ANTHROPIC_API_KEY).toBe("sk-ant-host-key");
|
||||
expect(result.checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(false);
|
||||
expect(result.checks.some((check) => check.code === "claude_acp_login_probe_unavailable")).toBe(false);
|
||||
// The lane still reports that API-key auth is in use.
|
||||
expect(result.checks.some((check) => check.code === "claude_acp_anthropic_api_key_detected")).toBe(true);
|
||||
// The host key value never enters a check.
|
||||
expect(JSON.stringify(result.checks)).not.toContain("sk-ant-host-key");
|
||||
});
|
||||
|
||||
it("runs the host login probe with the host CLAUDE_CODE_OAUTH_TOKEN on a local target", async () => {
|
||||
// A local ACP run inherits the host environment, so a host subscription
|
||||
// OAuth token authenticates the real run. The Test lane runs the login probe
|
||||
// with the same host token, so the probe env matches the credential the real
|
||||
// run receives. The probe then reports a real result, not a false
|
||||
// auth-required.
|
||||
process.env.CLAUDE_CODE_OAUTH_TOKEN = "oauth-host-token";
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp" },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
// The probe runs once with the host token, so it authenticates and reports
|
||||
// no false auth-required and no probe-unavailable check.
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
expect(spawnedEnv.CLAUDE_CODE_OAUTH_TOKEN).toBe("oauth-host-token");
|
||||
expect(result.checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(false);
|
||||
expect(result.checks.some((check) => check.code === "claude_acp_login_probe_unavailable")).toBe(false);
|
||||
// The lane reports that the configured OAuth token is in use.
|
||||
expect(result.checks.some((check) => check.code === "claude_oauth_token_configured")).toBe(true);
|
||||
// The host token value never enters a check.
|
||||
expect(JSON.stringify(result.checks)).not.toContain("oauth-host-token");
|
||||
});
|
||||
|
||||
it("runs the host login probe with the host ANTHROPIC_AUTH_TOKEN on a local target", async () => {
|
||||
// A local ACP run inherits the host environment, so a host bearer auth token
|
||||
// authenticates the real run. The Test lane runs the login probe with the
|
||||
// same host token, so the probe env matches the credential the real run
|
||||
// receives. The probe then reports a real result, not a false auth-required.
|
||||
process.env.ANTHROPIC_AUTH_TOKEN = "auth-host-token";
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp" },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
// The probe runs once with the host token, so it authenticates and reports
|
||||
// no false auth-required and no probe-unavailable check.
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
expect(spawnedEnv.ANTHROPIC_AUTH_TOKEN).toBe("auth-host-token");
|
||||
expect(result.checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(false);
|
||||
expect(result.checks.some((check) => check.code === "claude_acp_login_probe_unavailable")).toBe(false);
|
||||
// The host token value never enters a check.
|
||||
expect(JSON.stringify(result.checks)).not.toContain("auth-host-token");
|
||||
});
|
||||
|
||||
it("runs the host login probe with the host CLAUDE_CONFIG_DIR on a local target", async () => {
|
||||
// A local ACP run reads the stored Claude login from the host
|
||||
// CLAUDE_CONFIG_DIR. The Test lane runs the login probe with the same host
|
||||
// config dir, so the probe reads the same stored login the real run uses.
|
||||
// The probe then reports a real result, not a false auth-required.
|
||||
process.env.CLAUDE_CONFIG_DIR = "/host/claude/config";
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp" },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
// The probe runs once with the host config dir, so it reads the stored login
|
||||
// and reports no false auth-required and no probe-unavailable check.
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
expect(spawnedEnv.CLAUDE_CONFIG_DIR).toBe("/host/claude/config");
|
||||
expect(result.checks.some((check) => check.code === "claude_hello_probe_auth_required")).toBe(false);
|
||||
expect(result.checks.some((check) => check.code === "claude_acp_login_probe_unavailable")).toBe(false);
|
||||
});
|
||||
|
||||
it("never seeds the host ANTHROPIC_AUTH_TOKEN or CLAUDE_CONFIG_DIR on a remote target", async () => {
|
||||
// A remote target does not inherit the host environment, so the probe keeps
|
||||
// the deny-by-default env and never reads a host credential. The host token
|
||||
// and host config dir must never reach the remote probe env.
|
||||
process.env.ANTHROPIC_AUTH_TOKEN = "auth-host-token";
|
||||
process.env.CLAUDE_CONFIG_DIR = "/host/claude/config";
|
||||
probeResult.value = { exitCode: 0, stdout: helloStdout, stderr: "", timedOut: false };
|
||||
|
||||
await testClaudeAcpEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "acp" },
|
||||
executionTarget: sshTarget,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
expect(spawnedEnv.ANTHROPIC_AUTH_TOKEN).toBeUndefined();
|
||||
expect(spawnedEnv.CLAUDE_CONFIG_DIR).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -420,6 +420,11 @@ describe("claude_local ACP lane", () => {
|
|||
await fs.writeFile(commandPath, "#!/usr/bin/env sh\n", "utf8");
|
||||
setNodeVersion("v24.11.0");
|
||||
|
||||
// The ACP lane now verifies auth: without a credential the lane runs a real
|
||||
// host login probe, so its status depends on the host login state. Give the
|
||||
// config a Bedrock credential to make the auth path deterministic. Bedrock
|
||||
// gates off the host login probe, so the result reflects only the ACP
|
||||
// prerequisites, and a valid credential lets the lane report a pass.
|
||||
const result = await testClaudeAcpEnvironment({
|
||||
adapterType: "claude_local",
|
||||
companyId: "company-1",
|
||||
|
|
@ -427,6 +432,7 @@ describe("claude_local ACP lane", () => {
|
|||
engine: "acp",
|
||||
cwd: root,
|
||||
agentCommand: commandPath,
|
||||
env: { CLAUDE_CODE_USE_BEDROCK: "1" },
|
||||
},
|
||||
});
|
||||
|
||||
|
|
@ -443,6 +449,12 @@ describe("claude_local ACP lane", () => {
|
|||
level: "info",
|
||||
}),
|
||||
);
|
||||
expect(result.checks).toContainEqual(
|
||||
expect.objectContaining({
|
||||
code: "claude_acp_bedrock_auth",
|
||||
level: "info",
|
||||
}),
|
||||
);
|
||||
expect(result.checks).toContainEqual(
|
||||
expect.objectContaining({
|
||||
code: "claude_acp_runtime_scaffold",
|
||||
|
|
|
|||
|
|
@ -43,9 +43,12 @@ import {
|
|||
prepareSandboxClaudeProbeRuntime,
|
||||
} from "./claude-config.js";
|
||||
import {
|
||||
buildAdapterTestTargetCheck,
|
||||
buildClaudeLoginRequiredHint,
|
||||
logRedactedSandboxProbeDiagnostic,
|
||||
classifyThrownErrorClass,
|
||||
logSandboxProbeDiagnostic,
|
||||
} from "./probe-diagnostics.js";
|
||||
import { buildLocalAdapterTestProbeEnv } from "./probe-env.js";
|
||||
import { detectClaudeLoginRequired, parseClaudeStreamJson } from "./parse.js";
|
||||
import { buildClaudeProbePermissionArgs } from "./permissions.js";
|
||||
import { ADAPTER_AUTH_MISSING_CHECK_CODE } from "./auth-check.js";
|
||||
|
|
@ -494,27 +497,34 @@ function isNonEmpty(value: unknown): value is string {
|
|||
}
|
||||
|
||||
/**
|
||||
* Build the two checks that tell the user interface a sandbox has no ready
|
||||
* Claude authentication. The first check is descriptive for diagnostics. The
|
||||
* second check is the neutral canonical code the user interface reads to offer
|
||||
* login. The user interface does not read the message text or the top-level
|
||||
* status.
|
||||
* Build the checks that tell the user the probed target has no ready Claude
|
||||
* authentication. Every target gets the descriptive warn check, so
|
||||
* `summarizeStatus` never reports a pass without auth. Only a sandbox target
|
||||
* gets the neutral canonical `adapter_auth_missing` code, because only a
|
||||
* sandbox target can start an in-place login. The user interface reads the
|
||||
* canonical code to offer login and gates that affordance to sandbox targets.
|
||||
*/
|
||||
function buildAcpSandboxAuthMissingChecks(loginUrl: string | null): AdapterEnvironmentCheck[] {
|
||||
return [
|
||||
function buildAcpAuthMissingChecks(input: {
|
||||
targetIsSandbox: boolean;
|
||||
loginUrl: string | null;
|
||||
}): AdapterEnvironmentCheck[] {
|
||||
const checks: AdapterEnvironmentCheck[] = [
|
||||
{
|
||||
code: "claude_hello_probe_auth_required",
|
||||
level: "warn",
|
||||
message: "Claude ACP is available, but login is required.",
|
||||
hint: buildClaudeLoginRequiredHint(loginUrl),
|
||||
hint: buildClaudeLoginRequiredHint(input.loginUrl),
|
||||
},
|
||||
{
|
||||
];
|
||||
if (input.targetIsSandbox) {
|
||||
checks.push({
|
||||
code: ADAPTER_AUTH_MISSING_CHECK_CODE,
|
||||
level: "warn",
|
||||
message: "The sandbox has no ready authentication for this adapter.",
|
||||
hint: "Provide credentials for this adapter, or start login in the sandbox.",
|
||||
},
|
||||
];
|
||||
});
|
||||
}
|
||||
return checks;
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -522,64 +532,101 @@ function buildAcpSandboxAuthMissingChecks(loginUrl: string | null): AdapterEnvir
|
|||
* ACP path. The check is a warn, not an info, so `summarizeStatus` never
|
||||
* reports a pass. The check code is distinct from `adapter_auth_missing`, so the
|
||||
* user interface never shows the login affordance for a probe that could not
|
||||
* confirm the login state. Nicky's direction: a sandbox Test without available
|
||||
* auth must not report a success.
|
||||
* confirm the login state. A Test without available auth must not report a
|
||||
* success.
|
||||
*/
|
||||
function buildAcpLoginProbeUnavailableCheck(message: string): AdapterEnvironmentCheck {
|
||||
function buildAcpLoginProbeUnavailableCheck(
|
||||
message: string,
|
||||
targetIsSandbox = false,
|
||||
): AdapterEnvironmentCheck {
|
||||
return {
|
||||
code: "claude_acp_login_probe_unavailable",
|
||||
level: "warn",
|
||||
message,
|
||||
hint: "Verify that the sandbox can run `claude` and retry the Test. Set engine=cli to use the Claude CLI lane.",
|
||||
hint: targetIsSandbox
|
||||
? "Verify that the sandbox can run `claude` and retry the Test. Set engine=cli to use the Claude CLI lane."
|
||||
: "Verify that `claude` can run in this environment and retry the Test. Set engine=cli to use the Claude CLI lane.",
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Probe the stored Claude login inside a sandbox on the ACP path. The ACP engine
|
||||
* and the Claude CLI share the same stored Claude login, so the probe runs the
|
||||
* `claude` command with a short hello turn. The caller passes the prepared
|
||||
* `env`, so the probe reads the managed `CLAUDE_CONFIG_DIR` the same way the CLI
|
||||
* lane does. When the probe reports that login is required, the function returns
|
||||
* the canonical auth-missing checks. The user interface reads the canonical
|
||||
* check to offer login on the default ACP path, the same way it does for the
|
||||
* Claude CLI path.
|
||||
* Probe the stored Claude login for the probed target on the ACP path. The ACP
|
||||
* engine and the Claude CLI share the same stored Claude login, so the probe
|
||||
* runs the `claude` command with a short hello turn. The probe runs against any
|
||||
* target: a local host, an SSH remote, or a sandbox. On a local target the
|
||||
* probe builds the child env and the executable from the shared
|
||||
* deny-by-default builder, so a hostile caller value can neither select the
|
||||
* executable nor reach the child. On a remote target the caller passes the
|
||||
* prepared `env`, so the probe reads the managed `CLAUDE_CONFIG_DIR` the same
|
||||
* way the CLI lane does.
|
||||
*
|
||||
* The function keeps two signals distinct. It returns `adapter_auth_missing`
|
||||
* When the probe reports that login is required, the function returns the
|
||||
* auth-required checks. Only a sandbox target also gets the canonical
|
||||
* `adapter_auth_missing` code, so the user interface offers login for sandbox
|
||||
* targets only.
|
||||
*
|
||||
* The function keeps two signals distinct. It returns the auth-required check
|
||||
* only when the probe ran and login is required. It returns a separate warn
|
||||
* check when the probe could not run, timed out, or did not complete. It never
|
||||
* maps "probe could not run" to a silent pass.
|
||||
*/
|
||||
export async function probeClaudeAcpSandboxLogin(input: {
|
||||
config: Record<string, unknown>;
|
||||
target: AdapterExecutionTarget;
|
||||
target: AdapterExecutionTarget | null;
|
||||
env?: Record<string, string>;
|
||||
}): Promise<AdapterEnvironmentCheck[]> {
|
||||
const { config, target } = input;
|
||||
let env: Record<string, string>;
|
||||
const targetIsRemote = target?.kind === "remote";
|
||||
const targetIsSandbox = target?.kind === "remote" && target.transport === "sandbox";
|
||||
|
||||
// The caller-derived env. On a local target the shared builder filters it to
|
||||
// a deny-by-default allowlist. On a remote target the prepared env is used
|
||||
// directly, because the remote transport owns its own env sanitization.
|
||||
let callerEnv: Record<string, string>;
|
||||
if (input.env) {
|
||||
env = input.env;
|
||||
callerEnv = input.env;
|
||||
} else {
|
||||
const envConfig = parseObject(config.env);
|
||||
env = {};
|
||||
callerEnv = {};
|
||||
for (const [key, value] of Object.entries(envConfig)) {
|
||||
if (typeof value === "string") env[key] = value;
|
||||
if (typeof value === "string") callerEnv[key] = value;
|
||||
}
|
||||
}
|
||||
const command = "claude";
|
||||
|
||||
let command: string;
|
||||
let env: Record<string, string>;
|
||||
let cwd: string;
|
||||
if (targetIsRemote && target) {
|
||||
command = "claude";
|
||||
env = callerEnv;
|
||||
cwd = target.kind === "remote" ? target.remoteCwd : process.cwd();
|
||||
} else {
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv,
|
||||
trustedEnv: process.env,
|
||||
});
|
||||
if (!built.command) {
|
||||
return [buildAcpLoginProbeUnavailableCheck("Claude is not installed on the Paperclip host.")];
|
||||
}
|
||||
command = built.command;
|
||||
env = built.env;
|
||||
cwd = asString(config.cwd, process.cwd());
|
||||
}
|
||||
|
||||
const args = ["--print", "-", "--output-format", "stream-json", "--verbose"];
|
||||
args.push(
|
||||
...buildClaudeProbePermissionArgs({
|
||||
dangerouslySkipPermissions: asBoolean(config.dangerouslySkipPermissions, true),
|
||||
targetIsRemote: true,
|
||||
targetIsRemote,
|
||||
localProcessUid: process.getuid?.() ?? null,
|
||||
}),
|
||||
);
|
||||
const timeoutSec = Math.max(1, asNumber(config.helloProbeTimeoutSec, 90));
|
||||
const timeoutSec = Math.max(1, asNumber(config.helloProbeTimeoutSec, targetIsSandbox ? 90 : 45));
|
||||
const runId = `claude-acp-authprobe-${Date.now()}-${Math.random().toString(16).slice(2)}`;
|
||||
let probe: Awaited<ReturnType<typeof runAdapterExecutionTargetProcess>>;
|
||||
try {
|
||||
probe = await runAdapterExecutionTargetProcess(runId, target, command, args, {
|
||||
cwd: target.kind === "remote" ? target.remoteCwd : process.cwd(),
|
||||
cwd,
|
||||
env,
|
||||
timeoutSec,
|
||||
graceSec: 5,
|
||||
|
|
@ -587,16 +634,23 @@ export async function probeClaudeAcpSandboxLogin(input: {
|
|||
onLog: async () => {},
|
||||
});
|
||||
} catch (err) {
|
||||
// Keep the raw error out of the Test-result check. Send the redacted
|
||||
// diagnostic to the server log instead.
|
||||
logRedactedSandboxProbeDiagnostic(
|
||||
"Claude ACP login probe could not run in the sandbox",
|
||||
err instanceof Error ? err.message : String(err),
|
||||
);
|
||||
return [buildAcpLoginProbeUnavailableCheck("The Claude login probe could not run in the sandbox.")];
|
||||
// Keep the raw error out of the Test-result check and the server log. Log
|
||||
// only the fixed context, the allowlisted classification, and a safe error
|
||||
// class name.
|
||||
logSandboxProbeDiagnostic("Claude ACP login probe could not run", "spawn_error", {
|
||||
errorClass: classifyThrownErrorClass(err),
|
||||
});
|
||||
return [
|
||||
buildAcpLoginProbeUnavailableCheck(
|
||||
targetIsSandbox
|
||||
? "The Claude login probe could not run in the sandbox."
|
||||
: "The Claude login probe could not run.",
|
||||
targetIsSandbox,
|
||||
),
|
||||
];
|
||||
}
|
||||
if (probe.timedOut) {
|
||||
return [buildAcpLoginProbeUnavailableCheck("The Claude login probe timed out.")];
|
||||
return [buildAcpLoginProbeUnavailableCheck("The Claude login probe timed out.", targetIsSandbox)];
|
||||
}
|
||||
const parsedStream = parseClaudeStreamJson(probe.stdout);
|
||||
const loginMeta = detectClaudeLoginRequired({
|
||||
|
|
@ -605,16 +659,16 @@ export async function probeClaudeAcpSandboxLogin(input: {
|
|||
stderr: probe.stderr,
|
||||
});
|
||||
if (loginMeta.requiresLogin) {
|
||||
return buildAcpSandboxAuthMissingChecks(loginMeta.loginUrl);
|
||||
return buildAcpAuthMissingChecks({ targetIsSandbox, loginUrl: loginMeta.loginUrl });
|
||||
}
|
||||
if ((probe.exitCode ?? 1) !== 0) {
|
||||
// Keep the raw sandbox stderr and stdout out of the Test-result check. Send
|
||||
// the redacted diagnostic to the server log instead.
|
||||
logRedactedSandboxProbeDiagnostic(
|
||||
"Claude ACP login probe did not complete",
|
||||
firstNonEmptyString(probe.stderr, probe.stdout),
|
||||
);
|
||||
return [buildAcpLoginProbeUnavailableCheck("The Claude login probe did not complete.")];
|
||||
// Keep the raw stderr and stdout out of the Test-result check and the
|
||||
// server log. Log only the fixed context, the allowlisted classification,
|
||||
// and the safe exit code.
|
||||
logSandboxProbeDiagnostic("Claude ACP login probe did not complete", "nonzero_exit", {
|
||||
exitCode: probe.exitCode ?? null,
|
||||
});
|
||||
return [buildAcpLoginProbeUnavailableCheck("The Claude login probe did not complete.", targetIsSandbox)];
|
||||
}
|
||||
return [];
|
||||
}
|
||||
|
|
@ -626,6 +680,7 @@ export async function testClaudeAcpEnvironment(
|
|||
const config = parseObject(ctx.config);
|
||||
const target = ctx.executionTarget ?? null;
|
||||
const targetIsRemote = target?.kind === "remote";
|
||||
const targetIsSandbox = target?.kind === "remote" && target.transport === "sandbox";
|
||||
|
||||
checks.push({
|
||||
code: "claude_engine_selected",
|
||||
|
|
@ -634,6 +689,12 @@ export async function testClaudeAcpEnvironment(
|
|||
hint: "Set engine=cli to use the existing Claude Code CLI lane.",
|
||||
});
|
||||
|
||||
// Always name the target the Test probed, so a pass result never hides which
|
||||
// target it checked. A local probe reports the fixed host label.
|
||||
checks.push(
|
||||
buildAdapterTestTargetCheck({ targetIsRemote, environmentName: ctx.environmentName }),
|
||||
);
|
||||
|
||||
if (targetIsRemote) {
|
||||
checks.push({
|
||||
code: "claude_acp_remote_target",
|
||||
|
|
@ -698,6 +759,9 @@ export async function testClaudeAcpEnvironment(
|
|||
(considerHostEnv && isNonEmpty(process.env.ANTHROPIC_BEDROCK_BASE_URL));
|
||||
const configApiKey = envConfig.ANTHROPIC_API_KEY;
|
||||
const hostApiKey = considerHostEnv ? process.env.ANTHROPIC_API_KEY : undefined;
|
||||
const hostOauthToken = considerHostEnv ? process.env.CLAUDE_CODE_OAUTH_TOKEN : undefined;
|
||||
const hostAuthToken = considerHostEnv ? process.env.ANTHROPIC_AUTH_TOKEN : undefined;
|
||||
const hostConfigDir = considerHostEnv ? process.env.CLAUDE_CONFIG_DIR : undefined;
|
||||
if (hasBedrock) {
|
||||
checks.push({
|
||||
code: "claude_acp_bedrock_auth",
|
||||
|
|
@ -714,6 +778,20 @@ export async function testClaudeAcpEnvironment(
|
|||
detail: `Detected in ${source}.`,
|
||||
hint: "Unset ANTHROPIC_API_KEY if you want subscription-based Claude login behavior.",
|
||||
});
|
||||
} else if (
|
||||
isNonEmpty(envConfig.CLAUDE_CODE_OAUTH_TOKEN) ||
|
||||
(considerHostEnv && isNonEmpty(process.env.CLAUDE_CODE_OAUTH_TOKEN))
|
||||
) {
|
||||
const source = isNonEmpty(envConfig.CLAUDE_CODE_OAUTH_TOKEN)
|
||||
? "configured environment variables"
|
||||
: "server environment";
|
||||
checks.push({
|
||||
code: "claude_oauth_token_configured",
|
||||
level: "info",
|
||||
message:
|
||||
"CLAUDE_CODE_OAUTH_TOKEN is set. Claude ACP will authenticate with the configured subscription token; no stored login is needed on the execution target.",
|
||||
detail: `Detected in ${source}.`,
|
||||
});
|
||||
} else if (!targetIsRemote) {
|
||||
checks.push({
|
||||
code: "claude_acp_subscription_mode_possible",
|
||||
|
|
@ -722,24 +800,52 @@ export async function testClaudeAcpEnvironment(
|
|||
});
|
||||
}
|
||||
|
||||
// A sandbox target can start a login flow, and subscription auth is the only
|
||||
// credential source left after the branches above rule out Bedrock and an
|
||||
// API key. Prepare the sandbox the same way the CLI lane does — install the
|
||||
// Claude CLI when it is absent and materialize the managed CLAUDE_CONFIG_DIR
|
||||
// — then probe the stored Claude login. The Test result carries the canonical
|
||||
// adapter_auth_missing signal when login is required, and a distinct warn
|
||||
// check when the probe cannot run. The user interface reads the canonical
|
||||
// signal to offer login on the default ACP path.
|
||||
if (
|
||||
target?.kind === "remote" &&
|
||||
target.transport === "sandbox" &&
|
||||
!hasBedrock &&
|
||||
!isNonEmpty(configApiKey)
|
||||
) {
|
||||
// Run a real hello probe for every target when Bedrock and a config API key
|
||||
// are both absent. A local target inherits the host environment, so the real
|
||||
// ACP run authenticates with a host ANTHROPIC_API_KEY. The probe seeds the
|
||||
// same host key below when the config sets none, so the probe uses the
|
||||
// credential the real run receives and does not report a false auth-required.
|
||||
// A remote target does not inherit the host env, so considerHostEnv is false
|
||||
// and the probe never reads the host key. The CLI lane already probes every
|
||||
// target; the ACP lane now matches it, so a local or SSH target no longer
|
||||
// reports a pass without a credential check. Prepare the sandbox the same way
|
||||
// the CLI lane does — install the Claude CLI when it is absent and materialize
|
||||
// the managed CLAUDE_CONFIG_DIR. The preparation is a no-op for a local or SSH
|
||||
// target. The probe returns the canonical adapter_auth_missing signal only for
|
||||
// a sandbox target, and a distinct warn check when the probe cannot run. The
|
||||
// user interface reads the canonical signal to offer login on the sandbox ACP
|
||||
// path.
|
||||
if (!hasBedrock && !isNonEmpty(configApiKey)) {
|
||||
const probeEnv: Record<string, string> = {};
|
||||
for (const [key, value] of Object.entries(envConfig)) {
|
||||
if (typeof value === "string") probeEnv[key] = value;
|
||||
}
|
||||
// Seed the host ANTHROPIC_API_KEY when the config sets no key, so the probe
|
||||
// env matches the credential the real local run inherits from the host.
|
||||
if (isNonEmpty(hostApiKey) && !isNonEmpty(probeEnv.ANTHROPIC_API_KEY)) {
|
||||
probeEnv.ANTHROPIC_API_KEY = hostApiKey.trim();
|
||||
}
|
||||
// Seed the host CLAUDE_CODE_OAUTH_TOKEN the same way. A local ACP run
|
||||
// inherits a host subscription OAuth token, so the probe must receive the
|
||||
// same token. Without this seed a valid host OAuth-token setup reports a
|
||||
// false claude_hello_probe_auth_required and fails the Test lane.
|
||||
if (isNonEmpty(hostOauthToken) && !isNonEmpty(probeEnv.CLAUDE_CODE_OAUTH_TOKEN)) {
|
||||
probeEnv.CLAUDE_CODE_OAUTH_TOKEN = hostOauthToken.trim();
|
||||
}
|
||||
// Seed the host ANTHROPIC_AUTH_TOKEN the same way. A local ACP run inherits
|
||||
// a host bearer auth token, so the probe must receive the same token.
|
||||
// Without this seed a valid host ANTHROPIC_AUTH_TOKEN setup reports a false
|
||||
// claude_hello_probe_auth_required and fails the Test lane.
|
||||
if (isNonEmpty(hostAuthToken) && !isNonEmpty(probeEnv.ANTHROPIC_AUTH_TOKEN)) {
|
||||
probeEnv.ANTHROPIC_AUTH_TOKEN = hostAuthToken.trim();
|
||||
}
|
||||
// Seed the host CLAUDE_CONFIG_DIR the same way. A local ACP run reads the
|
||||
// stored Claude login from the host CLAUDE_CONFIG_DIR, so the probe must
|
||||
// read the same stored login. Without this seed a valid host stored login
|
||||
// reports a false claude_hello_probe_auth_required and fails the Test lane.
|
||||
if (isNonEmpty(hostConfigDir) && !isNonEmpty(probeEnv.CLAUDE_CONFIG_DIR)) {
|
||||
probeEnv.CLAUDE_CONFIG_DIR = hostConfigDir.trim();
|
||||
}
|
||||
const runId = `claude-acp-envtest-${Date.now()}-${Math.random().toString(16).slice(2)}`;
|
||||
checks.push(
|
||||
...(await prepareSandboxClaudeProbeRuntime({
|
||||
|
|
@ -750,9 +856,9 @@ export async function testClaudeAcpEnvironment(
|
|||
env: probeEnv,
|
||||
installCommand: SANDBOX_INSTALL_COMMAND,
|
||||
detectCommand: "claude",
|
||||
targetIsRemote: true,
|
||||
targetIsSandbox: true,
|
||||
helloProbeTimeoutSec: asNumber(config.helloProbeTimeoutSec, 90),
|
||||
targetIsRemote,
|
||||
targetIsSandbox,
|
||||
helloProbeTimeoutSec: asNumber(config.helloProbeTimeoutSec, targetIsSandbox ? 90 : 45),
|
||||
})),
|
||||
);
|
||||
const canProbe = !checks.some((check) => check.code === "claude_managed_config_dir_failed");
|
||||
|
|
|
|||
|
|
@ -2,7 +2,27 @@ import * as fs from "node:fs/promises";
|
|||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import { afterEach, describe, expect, it, vi } from "vitest";
|
||||
import { prepareClaudeConfigSeed } from "./claude-config.js";
|
||||
import type { AdapterExecutionTarget } from "@paperclipai/adapter-utils/execution-target";
|
||||
|
||||
// A shared handle so the managed-config test can force the runtime preparation
|
||||
// step to throw an error that carries untrusted markers.
|
||||
const { prepareAdapterExecutionTargetRuntime } = vi.hoisted(() => ({
|
||||
prepareAdapterExecutionTargetRuntime: vi.fn(),
|
||||
}));
|
||||
|
||||
vi.mock("@paperclipai/adapter-utils/execution-target", async () => {
|
||||
const actual = await vi.importActual<typeof import("@paperclipai/adapter-utils/execution-target")>(
|
||||
"@paperclipai/adapter-utils/execution-target",
|
||||
);
|
||||
return {
|
||||
...actual,
|
||||
adapterExecutionTargetUsesManagedHome: () => true,
|
||||
maybeRunSandboxInstallCommand: async () => null,
|
||||
prepareAdapterExecutionTargetRuntime,
|
||||
};
|
||||
});
|
||||
|
||||
import { prepareClaudeConfigSeed, prepareSandboxClaudeProbeRuntime } from "./claude-config.js";
|
||||
|
||||
describe("prepareClaudeConfigSeed", () => {
|
||||
const cleanupDirs: string[] = [];
|
||||
|
|
@ -109,3 +129,96 @@ describe("prepareClaudeConfigSeed", () => {
|
|||
.resolves.toBe("local instructions");
|
||||
});
|
||||
});
|
||||
|
||||
describe("prepareSandboxClaudeProbeRuntime managed-config diagnostics", () => {
|
||||
const cleanupDirs: string[] = [];
|
||||
const savedEnv: Record<string, string | undefined> = {};
|
||||
|
||||
const sandboxTarget: AdapterExecutionTarget = {
|
||||
kind: "remote",
|
||||
transport: "sandbox",
|
||||
providerKey: "daytona",
|
||||
remoteCwd: "/home/daytona/paperclip-workspace",
|
||||
runner: {
|
||||
execute: async () => ({
|
||||
exitCode: 0,
|
||||
signal: null,
|
||||
timedOut: false,
|
||||
stdout: "",
|
||||
stderr: "",
|
||||
pid: null,
|
||||
startedAt: new Date().toISOString(),
|
||||
}),
|
||||
},
|
||||
};
|
||||
|
||||
afterEach(async () => {
|
||||
vi.restoreAllMocks();
|
||||
vi.clearAllMocks();
|
||||
for (const [key, value] of Object.entries(savedEnv)) {
|
||||
if (value === undefined) delete process.env[key];
|
||||
else process.env[key] = value;
|
||||
}
|
||||
while (cleanupDirs.length > 0) {
|
||||
const dir = cleanupDirs.pop();
|
||||
if (!dir) continue;
|
||||
await fs.rm(dir, { recursive: true, force: true }).catch(() => undefined);
|
||||
}
|
||||
});
|
||||
|
||||
it("keeps a thrown config-materialization error out of every check and the log", async () => {
|
||||
// The runtime preparation throws an error that carries two untrusted values:
|
||||
// an opaque credential marker and a proxy marker. Neither may reach a check
|
||||
// or the server log. The log carries only the fixed context, the allowlisted
|
||||
// classification, and the safe error class name.
|
||||
const opaqueCredMarker = "OPAQUECREDMARKERconfig";
|
||||
const proxyMarker = "http://user:pass@proxy.corp.internal:3128";
|
||||
|
||||
const root = await fs.mkdtemp(path.join(os.tmpdir(), "paperclip-claude-config-mgmt-"));
|
||||
cleanupDirs.push(root);
|
||||
const sourceDir = path.join(root, "claude-source");
|
||||
await fs.mkdir(sourceDir, { recursive: true });
|
||||
|
||||
for (const key of ["CLAUDE_CONFIG_DIR", "PAPERCLIP_HOME", "PAPERCLIP_INSTANCE_ID"]) {
|
||||
savedEnv[key] = process.env[key];
|
||||
}
|
||||
process.env.CLAUDE_CONFIG_DIR = sourceDir;
|
||||
process.env.PAPERCLIP_HOME = path.join(root, "paperclip-home");
|
||||
process.env.PAPERCLIP_INSTANCE_ID = "test-instance";
|
||||
|
||||
prepareAdapterExecutionTargetRuntime.mockRejectedValueOnce(
|
||||
new Error(`materialize failed with ${opaqueCredMarker} via ${proxyMarker}`),
|
||||
);
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
|
||||
const checks = await prepareSandboxClaudeProbeRuntime({
|
||||
runId: "run-1",
|
||||
target: sandboxTarget,
|
||||
// The probe passes no CLAUDE_CONFIG_DIR, so the managed branch runs.
|
||||
cwd: "/home/daytona/paperclip-workspace",
|
||||
companyId: "company-1",
|
||||
env: {},
|
||||
installCommand: "install-claude",
|
||||
detectCommand: "claude",
|
||||
targetIsRemote: true,
|
||||
targetIsSandbox: true,
|
||||
helloProbeTimeoutSec: 30,
|
||||
});
|
||||
|
||||
const failed = checks.find((check) => check.code === "claude_managed_config_dir_failed");
|
||||
expect(failed).toBeTruthy();
|
||||
const checkText = JSON.stringify(checks);
|
||||
expect(checkText).not.toContain(opaqueCredMarker);
|
||||
expect(checkText).not.toContain("proxy.corp.internal");
|
||||
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(opaqueCredMarker);
|
||||
expect(loggedText).not.toContain("proxy.corp.internal");
|
||||
expect(warnSpy.mock.calls[0]?.[1]).toMatchObject({
|
||||
classification: "spawn_error",
|
||||
errorClass: "Error",
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -17,7 +17,7 @@ import {
|
|||
} from "@paperclipai/adapter-utils/execution-target";
|
||||
import { resolvePaperclipInstanceRootForAdapter } from "@paperclipai/adapter-utils/server-utils";
|
||||
import { shellQuote } from "@paperclipai/adapter-utils/ssh";
|
||||
import { logRedactedSandboxProbeDiagnostic } from "./probe-diagnostics.js";
|
||||
import { classifyThrownErrorClass, logSandboxProbeDiagnostic } from "./probe-diagnostics.js";
|
||||
|
||||
const SEEDED_SHARED_FILES = ["settings.json", "CLAUDE.md"] as const;
|
||||
|
||||
|
|
@ -349,11 +349,13 @@ export async function prepareSandboxClaudeProbeRuntime(input: {
|
|||
detail: remoteClaudeConfigDir,
|
||||
});
|
||||
} catch (err) {
|
||||
// Keep the raw error out of the Test-result check. Send the redacted
|
||||
// diagnostic to the server log instead.
|
||||
logRedactedSandboxProbeDiagnostic(
|
||||
// Keep the raw error out of the Test-result check and the server log. Log
|
||||
// only the fixed context, the allowlisted classification, and a safe
|
||||
// error class name.
|
||||
logSandboxProbeDiagnostic(
|
||||
"Could not materialize Paperclip-managed Claude config for the sandbox probe",
|
||||
err instanceof Error ? err.message : String(err),
|
||||
"spawn_error",
|
||||
{ errorClass: classifyThrownErrorClass(err) },
|
||||
);
|
||||
checks.push({
|
||||
code: "claude_managed_config_dir_failed",
|
||||
|
|
|
|||
|
|
@ -1,7 +1,8 @@
|
|||
import { describe, expect, it, vi } from "vitest";
|
||||
import {
|
||||
buildClaudeLoginRequiredHint,
|
||||
logRedactedSandboxProbeDiagnostic,
|
||||
classifyThrownErrorClass,
|
||||
logSandboxProbeDiagnostic,
|
||||
normalizeClaudeLoginUrl,
|
||||
} from "./probe-diagnostics.js";
|
||||
|
||||
|
|
@ -69,41 +70,76 @@ describe("buildClaudeLoginRequiredHint", () => {
|
|||
});
|
||||
});
|
||||
|
||||
describe("logRedactedSandboxProbeDiagnostic", () => {
|
||||
it("redacts a JSON secret field before it reaches the log", () => {
|
||||
describe("classifyThrownErrorClass", () => {
|
||||
it("returns the constructor name for an Error", () => {
|
||||
expect(classifyThrownErrorClass(new TypeError("boom"))).toBe("TypeError");
|
||||
expect(classifyThrownErrorClass(new Error("boom"))).toBe("Error");
|
||||
});
|
||||
|
||||
it("returns null for a non-Error value", () => {
|
||||
expect(classifyThrownErrorClass("a raw secret string")).toBeNull();
|
||||
expect(classifyThrownErrorClass(null)).toBeNull();
|
||||
expect(classifyThrownErrorClass({ message: "opaque" })).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe("logSandboxProbeDiagnostic", () => {
|
||||
it("logs only the fixed context and the allowlisted classification", () => {
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
logRedactedSandboxProbeDiagnostic("probe failed", '{"token":"opaque-secret-value"}');
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain("opaque-secret-value");
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
logSandboxProbeDiagnostic("probe failed", "auth_required");
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
expect(warnSpy).toHaveBeenCalledWith("[paperclip] probe failed", {
|
||||
classification: "auth_required",
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("does not log when the diagnostic is empty", () => {
|
||||
it("adds a finite exit code as a structured field", () => {
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
logRedactedSandboxProbeDiagnostic("probe failed", "");
|
||||
expect(warnSpy).not.toHaveBeenCalled();
|
||||
logSandboxProbeDiagnostic("probe failed", "nonzero_exit", { exitCode: 3 });
|
||||
expect(warnSpy).toHaveBeenCalledWith("[paperclip] probe failed", {
|
||||
classification: "nonzero_exit",
|
||||
exitCode: 3,
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("redacts a JSON secret value with an escaped quote before it reaches the log", () => {
|
||||
it("drops a null or non-finite exit code", () => {
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
// The value holds an escaped quote, then the rest of the credential. A naive
|
||||
// matcher stops at the escaped quote and leaks the marker to the log.
|
||||
logRedactedSandboxProbeDiagnostic("probe failed", '{"token":"pre\\"MARKERLOGQUOTE"}');
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain("MARKERLOGQUOTE");
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
logSandboxProbeDiagnostic("probe failed", "nonzero_exit", { exitCode: null });
|
||||
logSandboxProbeDiagnostic("probe failed", "nonzero_exit", { exitCode: Number.NaN });
|
||||
for (const call of warnSpy.mock.calls) {
|
||||
expect(call[1]).toEqual({ classification: "nonzero_exit" });
|
||||
}
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("redacts an escaped-JSON secret value before it reaches the log", () => {
|
||||
it("sanitizes the error class to an identifier and bounds its length", () => {
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
const innerJson = '{"password":"pre\\\\MARKERLOGBACKSLASH"}';
|
||||
logRedactedSandboxProbeDiagnostic("probe failed", JSON.stringify(innerJson));
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain("MARKERLOGBACKSLASH");
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
// A crafted error class name that carries an opaque marker and a proxy URL.
|
||||
// The sanitizer must strip every non-identifier character and bound the
|
||||
// length, so no structured secret shape reaches the log.
|
||||
logSandboxProbeDiagnostic("probe failed", "spawn_error", {
|
||||
errorClass: `MARKER-LEAK http://user:pass@proxy.internal:8080/path?t=${"x".repeat(200)}`,
|
||||
});
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
const detail = warnSpy.mock.calls[0]![1] as { errorClass?: string };
|
||||
expect(detail.errorClass).toMatch(/^[A-Za-z0-9_$]+$/);
|
||||
expect(detail.errorClass!.length).toBeLessThanOrEqual(64);
|
||||
// The separators and the proxy structure do not survive.
|
||||
expect(detail.errorClass).not.toContain("-");
|
||||
expect(detail.errorClass).not.toContain(":");
|
||||
expect(detail.errorClass).not.toContain("/");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("drops an empty or non-string error class", () => {
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
logSandboxProbeDiagnostic("probe failed", "spawn_error", { errorClass: null });
|
||||
logSandboxProbeDiagnostic("probe failed", "spawn_error", { errorClass: "***" });
|
||||
for (const call of warnSpy.mock.calls) {
|
||||
expect(call[1]).toEqual({ classification: "spawn_error" });
|
||||
}
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,14 +1,74 @@
|
|||
import { redactDiagnosticText } from "@paperclipai/adapter-utils";
|
||||
import type { AdapterEnvironmentCheck } from "@paperclipai/adapter-utils";
|
||||
|
||||
// The server log keeps a bounded diagnostic. The bound stops a very large probe
|
||||
// output from filling the log.
|
||||
const MAX_LOGGED_PROBE_DIAGNOSTIC_CHARS = 2000;
|
||||
/**
|
||||
* The fixed label a Test result shows when the probe runs on the local
|
||||
* Paperclip host. The label is a constant, so a local target check never
|
||||
* carries an environment ID, a config value, or a credential-derived string.
|
||||
*/
|
||||
export const ADAPTER_TEST_HOST_TARGET_LABEL = "Paperclip host";
|
||||
|
||||
// The login hint may show a login URL. The URL must be a normalized https URL
|
||||
// with an allowlisted Claude or Anthropic host and no query or fragment. A host
|
||||
// matches when it equals a suffix or ends with a dot and the suffix.
|
||||
const ALLOWED_LOGIN_URL_HOST_SUFFIXES = ["anthropic.com", "claude.ai"] as const;
|
||||
|
||||
// A JavaScript error class name is a bounded identifier. The bound stops a very
|
||||
// large or crafted class name from filling the log.
|
||||
const MAX_ERROR_CLASS_NAME_CHARS = 64;
|
||||
|
||||
/**
|
||||
* The allowlisted classification for a sandbox probe diagnostic. The call site
|
||||
* picks one fixed label from this set. The label never holds a copy of
|
||||
* untrusted probe text.
|
||||
*
|
||||
* - `timeout`: the probe process did not finish before the deadline.
|
||||
* - `auth_required`: the probe ran and reported that login is required.
|
||||
* - `nonzero_exit`: the probe process exited with a non-zero exit code.
|
||||
* - `spawn_error`: the probe process, or a setup step, threw before it ran.
|
||||
* - `empty_output`: the probe produced no output.
|
||||
* - `unexpected_output`: the probe ran and exited zero, but the output did
|
||||
* not match the expected reply.
|
||||
*/
|
||||
export type SandboxProbeDiagnosticClassification =
|
||||
| "timeout"
|
||||
| "auth_required"
|
||||
| "nonzero_exit"
|
||||
| "spawn_error"
|
||||
| "empty_output"
|
||||
| "unexpected_output";
|
||||
|
||||
/**
|
||||
* The safe structured fields a call site may add to a probe diagnostic. Each
|
||||
* field is a fixed shape, not free text. The helper drops any value that is not
|
||||
* safe.
|
||||
*/
|
||||
export interface SandboxProbeDiagnosticFields {
|
||||
// The process exit code. The helper logs it only when it is a finite number.
|
||||
exitCode?: number | null;
|
||||
// The class name of a thrown value. Use `classifyThrownErrorClass` to derive
|
||||
// it. The helper sanitizes it again before it logs it.
|
||||
errorClass?: string | null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the class name of a thrown value for a safe probe diagnostic. The
|
||||
* function returns the constructor name of an `Error`, or `null` for any other
|
||||
* value. The name is a bounded identifier, not a copy of the error message, so
|
||||
* it carries no untrusted probe text.
|
||||
*/
|
||||
export function classifyThrownErrorClass(err: unknown): string | null {
|
||||
if (err instanceof Error) return err.constructor?.name ?? "Error";
|
||||
return null;
|
||||
}
|
||||
|
||||
// Keep only identifier characters and bound the length. The result cannot carry
|
||||
// untrusted probe text.
|
||||
function sanitizeErrorClassName(name: string | null | undefined): string | null {
|
||||
if (typeof name !== "string") return null;
|
||||
const safe = name.replace(/[^A-Za-z0-9_$]/g, "").slice(0, MAX_ERROR_CLASS_NAME_CHARS);
|
||||
return safe.length > 0 ? safe : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Send a sandbox probe or config materialization diagnostic to the server log.
|
||||
*
|
||||
|
|
@ -17,28 +77,35 @@ const ALLOWED_LOGIN_URL_HOST_SUFFIXES = ["anthropic.com", "claude.ai"] as const;
|
|||
* - the Claude ACP Test lane (`acp.ts`),
|
||||
* - the managed-config materialization step (`claude-config.ts`).
|
||||
*
|
||||
* The helper is the single boundary where a raw probe error, stdout, or stderr
|
||||
* string reaches an output. It redacts secrets with `redactDiagnosticText`
|
||||
* first, so no credential reaches the log. The sanitizer redacts shell
|
||||
* `KEY=value` secrets and JSON secret fields such as `{"token":"..."}`. The
|
||||
* helper also bounds the length. A caller must never copy the raw string into a
|
||||
* Test-result check, because the user interface renders check text.
|
||||
* Contract: no untrusted text reaches the log. The helper logs only the fixed
|
||||
* context string, one allowlisted classification, and safe structured fields.
|
||||
* The safe fields are the process exit code and a sanitized error class name.
|
||||
* The helper never logs raw probe stdout, raw stderr, or a raw thrown-error
|
||||
* message. A call site must never copy raw probe text into a Test-result check,
|
||||
* because the user interface renders check text.
|
||||
*
|
||||
* @param context A short fixed description of the failed step. It carries no
|
||||
* untrusted text.
|
||||
* @param raw The untrusted diagnostic from the sandbox. The helper redacts it.
|
||||
* @param classification One allowlisted label that the call site derives from
|
||||
* the probe state. It carries no untrusted text.
|
||||
* @param fields Optional safe structured fields, such as the exit code.
|
||||
*/
|
||||
export function logRedactedSandboxProbeDiagnostic(
|
||||
export function logSandboxProbeDiagnostic(
|
||||
context: string,
|
||||
raw: string | null | undefined,
|
||||
classification: SandboxProbeDiagnosticClassification,
|
||||
fields?: SandboxProbeDiagnosticFields,
|
||||
): void {
|
||||
if (!raw) return;
|
||||
const redacted = redactDiagnosticText(raw)
|
||||
.replace(/\s+/g, " ")
|
||||
.trim()
|
||||
.slice(0, MAX_LOGGED_PROBE_DIAGNOSTIC_CHARS);
|
||||
if (!redacted) return;
|
||||
console.warn(`[paperclip] ${context}`, { detail: redacted });
|
||||
const detail: {
|
||||
classification: SandboxProbeDiagnosticClassification;
|
||||
exitCode?: number;
|
||||
errorClass?: string;
|
||||
} = { classification };
|
||||
if (typeof fields?.exitCode === "number" && Number.isFinite(fields.exitCode)) {
|
||||
detail.exitCode = fields.exitCode;
|
||||
}
|
||||
const errorClass = sanitizeErrorClassName(fields?.errorClass);
|
||||
if (errorClass) detail.errorClass = errorClass;
|
||||
console.warn(`[paperclip] ${context}`, detail);
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -85,3 +152,39 @@ export function buildClaudeLoginRequiredHint(loginUrl: string | null | undefined
|
|||
? `Run \`claude login\` and complete sign-in at ${safeUrl}, then retry.`
|
||||
: "Run `claude login` in this environment, then retry the probe.";
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the label a Test result shows for the probed target. A remote target
|
||||
* uses the authorized environment name. A remote target with no name uses a
|
||||
* fixed generic label. A local target uses the fixed host label. The function
|
||||
* never returns an environment ID, a config value, or a credential-derived
|
||||
* string.
|
||||
*/
|
||||
export function resolveAdapterTestTargetLabel(input: {
|
||||
targetIsRemote: boolean;
|
||||
environmentName: string | null | undefined;
|
||||
}): string {
|
||||
if (!input.targetIsRemote) return ADAPTER_TEST_HOST_TARGET_LABEL;
|
||||
const name = typeof input.environmentName === "string" ? input.environmentName.trim() : "";
|
||||
return name.length > 0 ? name : "the selected environment";
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the target check every Test result carries, so the result names the
|
||||
* target it probed. Both the Claude CLI Test lane and the Claude ACP Test lane
|
||||
* use this builder. The check text carries only the authorized environment
|
||||
* label or the fixed host label.
|
||||
*/
|
||||
export function buildAdapterTestTargetCheck(input: {
|
||||
targetIsRemote: boolean;
|
||||
environmentName: string | null | undefined;
|
||||
}): AdapterEnvironmentCheck {
|
||||
const label = resolveAdapterTestTargetLabel(input);
|
||||
return {
|
||||
code: "claude_environment_target",
|
||||
level: "info",
|
||||
message: input.targetIsRemote
|
||||
? `Probing inside environment: ${label}`
|
||||
: "Probing on the Paperclip host.",
|
||||
};
|
||||
}
|
||||
|
|
|
|||
|
|
@ -0,0 +1,142 @@
|
|||
import { afterEach, describe, expect, it } from "vitest";
|
||||
import { chmod, mkdtemp, rm, writeFile } from "node:fs/promises";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import { buildLocalAdapterTestProbeEnv } from "./probe-env.js";
|
||||
|
||||
const tempDirs: string[] = [];
|
||||
|
||||
async function makeTrustedPathWithClaude(): Promise<{ dir: string; claudePath: string }> {
|
||||
const dir = await mkdtemp(path.join(os.tmpdir(), "paperclip-probe-env-"));
|
||||
tempDirs.push(dir);
|
||||
const claudePath = path.join(dir, "claude");
|
||||
await writeFile(claudePath, "#!/bin/sh\nexit 0\n");
|
||||
await chmod(claudePath, 0o755);
|
||||
return { dir, claudePath };
|
||||
}
|
||||
|
||||
afterEach(async () => {
|
||||
while (tempDirs.length > 0) {
|
||||
const dir = tempDirs.pop();
|
||||
if (dir) await rm(dir, { recursive: true, force: true }).catch(() => undefined);
|
||||
}
|
||||
});
|
||||
|
||||
describe("buildLocalAdapterTestProbeEnv", () => {
|
||||
it("resolves claude from the trusted PATH and ignores the caller PATH", async () => {
|
||||
const { dir, claudePath } = await makeTrustedPathWithClaude();
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: { PATH: "/hostile/bin", Path: "/hostile/bin", command: "/tmp/evil/claude" },
|
||||
trustedEnv: { PATH: dir },
|
||||
});
|
||||
expect(built.command).toBe(claudePath);
|
||||
expect(built.env.PATH).toBeUndefined();
|
||||
expect(built.env.Path).toBeUndefined();
|
||||
});
|
||||
|
||||
it("returns a null command when the trusted PATH holds no claude", async () => {
|
||||
const dir = await mkdtemp(path.join(os.tmpdir(), "paperclip-probe-env-empty-"));
|
||||
tempDirs.push(dir);
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: {},
|
||||
trustedEnv: { PATH: dir },
|
||||
});
|
||||
expect(built.command).toBeNull();
|
||||
});
|
||||
|
||||
it("keeps the allowlisted Claude, auth, and Bedrock values", async () => {
|
||||
const { dir } = await makeTrustedPathWithClaude();
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: {
|
||||
ANTHROPIC_API_KEY: "api-key-value",
|
||||
CLAUDE_CODE_OAUTH_TOKEN: "oauth-token-value",
|
||||
CLAUDE_CODE_USE_BEDROCK: "1",
|
||||
ANTHROPIC_BEDROCK_BASE_URL: "https://bedrock.example",
|
||||
AWS_ACCESS_KEY_ID: "aws-key",
|
||||
AWS_SECRET_ACCESS_KEY: "aws-secret",
|
||||
AWS_REGION: "us-east-1",
|
||||
CLAUDE_CONFIG_DIR: "/managed/config",
|
||||
},
|
||||
trustedEnv: { PATH: dir },
|
||||
});
|
||||
expect(built.env.ANTHROPIC_API_KEY).toBe("api-key-value");
|
||||
expect(built.env.CLAUDE_CODE_OAUTH_TOKEN).toBe("oauth-token-value");
|
||||
expect(built.env.CLAUDE_CODE_USE_BEDROCK).toBe("1");
|
||||
expect(built.env.ANTHROPIC_BEDROCK_BASE_URL).toBe("https://bedrock.example");
|
||||
expect(built.env.AWS_ACCESS_KEY_ID).toBe("aws-key");
|
||||
expect(built.env.AWS_SECRET_ACCESS_KEY).toBe("aws-secret");
|
||||
expect(built.env.AWS_REGION).toBe("us-east-1");
|
||||
expect(built.env.CLAUDE_CONFIG_DIR).toBe("/managed/config");
|
||||
});
|
||||
|
||||
it("drops hostile loader, PATH, shell-startup, and Windows interpreter keys", async () => {
|
||||
const { dir } = await makeTrustedPathWithClaude();
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: {
|
||||
PATH: "/hostile/bin",
|
||||
Path: "/hostile/bin",
|
||||
PATHEXT: ".EVIL",
|
||||
LD_PRELOAD: "/hostile/lib/evil.so",
|
||||
LD_LIBRARY_PATH: "/hostile/lib",
|
||||
DYLD_INSERT_LIBRARIES: "/hostile/lib/evil.dylib",
|
||||
DYLD_LIBRARY_PATH: "/hostile/lib",
|
||||
NODE_OPTIONS: "--require /hostile/evil.js",
|
||||
ENV: "/hostile/profile",
|
||||
BASH_ENV: "/hostile/bashrc",
|
||||
SystemRoot: "C:\\hostile",
|
||||
systemroot: "C:\\hostile",
|
||||
WINDIR: "C:\\hostile",
|
||||
windir: "C:\\hostile",
|
||||
ComSpec: "C:\\hostile\\evil.exe",
|
||||
comspec: "C:\\hostile\\evil.exe",
|
||||
},
|
||||
trustedEnv: { PATH: dir },
|
||||
});
|
||||
for (const key of Object.keys(built.env)) {
|
||||
expect(key.toUpperCase()).not.toBe("PATH");
|
||||
expect(key.toUpperCase()).not.toBe("PATHEXT");
|
||||
expect(key.toUpperCase()).not.toBe("LD_PRELOAD");
|
||||
expect(key.toUpperCase()).not.toBe("LD_LIBRARY_PATH");
|
||||
expect(key.toUpperCase()).not.toBe("DYLD_INSERT_LIBRARIES");
|
||||
expect(key.toUpperCase()).not.toBe("DYLD_LIBRARY_PATH");
|
||||
expect(key.toUpperCase()).not.toBe("NODE_OPTIONS");
|
||||
expect(key.toUpperCase()).not.toBe("ENV");
|
||||
expect(key.toUpperCase()).not.toBe("BASH_ENV");
|
||||
expect(key.toUpperCase()).not.toBe("SYSTEMROOT");
|
||||
expect(key.toUpperCase()).not.toBe("WINDIR");
|
||||
expect(key.toUpperCase()).not.toBe("COMSPEC");
|
||||
}
|
||||
});
|
||||
|
||||
it("takes proxy values only from the trusted env, never from the caller", async () => {
|
||||
const { dir } = await makeTrustedPathWithClaude();
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: {
|
||||
HTTP_PROXY: "http://caller-proxy:8080",
|
||||
HTTPS_PROXY: "http://caller-proxy:8443",
|
||||
http_proxy: "http://caller-proxy-lower:8080",
|
||||
NO_PROXY: "caller.example",
|
||||
},
|
||||
trustedEnv: { PATH: dir, HTTPS_PROXY: "http://trusted-proxy:8443" },
|
||||
});
|
||||
// The trusted proxy reaches the child; the caller proxy does not.
|
||||
expect(built.env.HTTPS_PROXY).toBe("http://trusted-proxy:8443");
|
||||
expect(built.env.HTTP_PROXY).toBeUndefined();
|
||||
expect(built.env.http_proxy).toBeUndefined();
|
||||
expect(built.env.NO_PROXY).toBeUndefined();
|
||||
// No env value carries the caller proxy host.
|
||||
const serialized = JSON.stringify(built.env);
|
||||
expect(serialized).not.toContain("caller-proxy");
|
||||
});
|
||||
|
||||
it("forwards no proxy variable when the trusted env has none", async () => {
|
||||
const { dir } = await makeTrustedPathWithClaude();
|
||||
const built = await buildLocalAdapterTestProbeEnv({
|
||||
callerEnv: { HTTP_PROXY: "http://caller-proxy:8080", https_proxy: "http://caller:8443" },
|
||||
trustedEnv: { PATH: dir },
|
||||
});
|
||||
for (const key of Object.keys(built.env)) {
|
||||
expect(key.toUpperCase()).not.toContain("PROXY");
|
||||
}
|
||||
});
|
||||
});
|
||||
|
|
@ -0,0 +1,187 @@
|
|||
import { access } from "node:fs/promises";
|
||||
import { constants as fsConstants } from "node:fs";
|
||||
import path from "node:path";
|
||||
|
||||
/**
|
||||
* The environment variable names that a local Claude adapter-test probe may
|
||||
* take from the untrusted adapter configuration. The builder denies every
|
||||
* other key by default. The list holds the documented Claude, Anthropic auth,
|
||||
* and AWS Bedrock variables that the probe needs to reach the real credential
|
||||
* the agent run uses.
|
||||
*/
|
||||
const LOCAL_PROBE_ALLOWED_CALLER_ENV_KEYS = [
|
||||
// Claude and Anthropic subscription and API auth.
|
||||
"ANTHROPIC_API_KEY",
|
||||
"ANTHROPIC_AUTH_TOKEN",
|
||||
"CLAUDE_CODE_OAUTH_TOKEN",
|
||||
"ANTHROPIC_BASE_URL",
|
||||
"ANTHROPIC_MODEL",
|
||||
"ANTHROPIC_SMALL_FAST_MODEL",
|
||||
"CLAUDE_CONFIG_DIR",
|
||||
// AWS Bedrock inference.
|
||||
"CLAUDE_CODE_USE_BEDROCK",
|
||||
"ANTHROPIC_BEDROCK_BASE_URL",
|
||||
"AWS_BEARER_TOKEN_BEDROCK",
|
||||
"AWS_ACCESS_KEY_ID",
|
||||
"AWS_SECRET_ACCESS_KEY",
|
||||
"AWS_SESSION_TOKEN",
|
||||
"AWS_PROFILE",
|
||||
"AWS_REGION",
|
||||
"AWS_DEFAULT_REGION",
|
||||
] as const;
|
||||
|
||||
/**
|
||||
* The proxy variable names the probe may forward. The builder reads these only
|
||||
* from the trusted server-resolved input. It never reads a proxy key from the
|
||||
* untrusted caller input, and it never logs, returns, or reflects a proxy
|
||||
* value.
|
||||
*/
|
||||
const PROXY_ENV_KEYS = ["HTTP_PROXY", "HTTPS_PROXY", "NO_PROXY"] as const;
|
||||
|
||||
/**
|
||||
* Windows interpreter-selection variables. The builder never takes these from
|
||||
* any input. The child-process launcher derives the interpreter from trusted
|
||||
* server state instead, so a caller value must never reach the child env.
|
||||
*/
|
||||
const WINDOWS_INTERPRETER_ENV_KEYS = new Set(["SYSTEMROOT", "WINDIR", "COMSPEC"]);
|
||||
|
||||
export interface LocalProbeEnvironment {
|
||||
/**
|
||||
* The trusted absolute path to the resolved `claude` executable, or `null`
|
||||
* when the trusted server PATH holds no `claude`. The caller must not run a
|
||||
* local probe when this is `null`; it reports a probe-unavailable check
|
||||
* instead.
|
||||
*/
|
||||
command: string | null;
|
||||
/**
|
||||
* The child environment for the local probe. It holds only allowlisted
|
||||
* caller values plus proxy values from the trusted input. It never holds a
|
||||
* caller-supplied proxy key, a loader variable, a PATH override, or a Windows
|
||||
* interpreter variable.
|
||||
*/
|
||||
env: Record<string, string>;
|
||||
}
|
||||
|
||||
function isNonEmptyString(value: unknown): value is string {
|
||||
return typeof value === "string" && value.trim().length > 0;
|
||||
}
|
||||
|
||||
function readCaseInsensitive(
|
||||
source: Record<string, string | undefined>,
|
||||
key: string,
|
||||
): string | undefined {
|
||||
const direct = source[key];
|
||||
if (typeof direct === "string") return direct;
|
||||
const upper = key.toUpperCase();
|
||||
for (const [candidateKey, candidateValue] of Object.entries(source)) {
|
||||
if (candidateKey.toUpperCase() === upper && typeof candidateValue === "string") {
|
||||
return candidateValue;
|
||||
}
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a command name to a trusted absolute executable. The function reads
|
||||
* the PATH and PATHEXT from the trusted server env only. It never reads a
|
||||
* caller value, so a hostile caller PATH cannot select the executable. The
|
||||
* function ignores a command that contains a path separator; a local probe
|
||||
* must not run a caller-supplied executable path.
|
||||
*/
|
||||
async function resolveTrustedExecutable(
|
||||
commandName: string,
|
||||
trustedEnv: NodeJS.ProcessEnv,
|
||||
): Promise<string | null> {
|
||||
if (commandName.includes("/") || commandName.includes("\\")) {
|
||||
return null;
|
||||
}
|
||||
const pathValue = trustedEnv.PATH ?? trustedEnv.Path ?? "";
|
||||
const delimiter = process.platform === "win32" ? ";" : ":";
|
||||
const dirs = pathValue.split(delimiter).filter(Boolean);
|
||||
const exts =
|
||||
process.platform === "win32"
|
||||
? (trustedEnv.PATHEXT ?? ".EXE;.CMD;.BAT;.COM").split(";").filter(Boolean)
|
||||
: [""];
|
||||
const mode = process.platform === "win32" ? fsConstants.F_OK : fsConstants.X_OK;
|
||||
for (const dir of dirs) {
|
||||
const candidates =
|
||||
process.platform === "win32"
|
||||
? exts.map((ext) => path.join(dir, `${commandName}${ext}`))
|
||||
: [path.join(dir, commandName)];
|
||||
for (const candidate of candidates) {
|
||||
try {
|
||||
await access(candidate, mode);
|
||||
return candidate;
|
||||
} catch {
|
||||
// Try the next candidate.
|
||||
}
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the child environment and executable for a local Claude adapter-test
|
||||
* hello probe. Both the ACP Test lane and the Claude CLI Test lane use this
|
||||
* builder, so the two lanes probe the host the same safe way.
|
||||
*
|
||||
* The builder denies by default. It never merges the arbitrary caller env into
|
||||
* a host base. It reads two separate inputs:
|
||||
* - `callerEnv`: the untrusted adapter-config env from the Test request. The
|
||||
* builder takes only allowlisted Claude, auth, and Bedrock keys from it. It
|
||||
* discards every proxy key and every Windows interpreter key, in any case.
|
||||
* - `trustedEnv`: the trusted server-resolved env (the server launch config
|
||||
* and the authorized environment's server-resolved env). The builder takes
|
||||
* proxy keys only from this input.
|
||||
*
|
||||
* The builder resolves `claude` to a trusted absolute executable with the
|
||||
* trusted server PATH, before it reads any caller value. It ignores a caller
|
||||
* `command` path.
|
||||
*
|
||||
* The builder never logs, returns, or reflects a proxy value in a check.
|
||||
*/
|
||||
export async function buildLocalAdapterTestProbeEnv(input: {
|
||||
callerEnv: Record<string, string>;
|
||||
trustedEnv?: NodeJS.ProcessEnv;
|
||||
commandName?: string;
|
||||
}): Promise<LocalProbeEnvironment> {
|
||||
const trustedEnv = input.trustedEnv ?? process.env;
|
||||
const commandName = input.commandName ?? "claude";
|
||||
const command = await resolveTrustedExecutable(commandName, trustedEnv);
|
||||
|
||||
const env: Record<string, string> = {};
|
||||
|
||||
// Allowlisted caller values. Read each allowed key by name so no unexpected
|
||||
// caller key can enter the child env. Proxy and Windows interpreter keys are
|
||||
// never in the allowlist, so a caller cannot pass them here.
|
||||
for (const key of LOCAL_PROBE_ALLOWED_CALLER_ENV_KEYS) {
|
||||
const value = readCaseInsensitive(input.callerEnv, key);
|
||||
if (isNonEmptyString(value)) {
|
||||
env[key] = value;
|
||||
}
|
||||
}
|
||||
|
||||
// Proxy values come from the trusted input only. A caller-supplied proxy key
|
||||
// in `callerEnv` is never read, so it cannot reach the child.
|
||||
for (const key of PROXY_ENV_KEYS) {
|
||||
const upperValue = trustedEnv[key];
|
||||
if (isNonEmptyString(upperValue)) {
|
||||
env[key] = upperValue;
|
||||
}
|
||||
const lowerKey = key.toLowerCase();
|
||||
const lowerValue = trustedEnv[lowerKey];
|
||||
if (isNonEmptyString(lowerValue)) {
|
||||
env[lowerKey] = lowerValue;
|
||||
}
|
||||
}
|
||||
|
||||
// Defense in depth: strip any Windows interpreter key that a future allowlist
|
||||
// edit might introduce, in any case.
|
||||
for (const key of Object.keys(env)) {
|
||||
if (WINDOWS_INTERPRETER_ENV_KEYS.has(key.toUpperCase())) {
|
||||
delete env[key];
|
||||
}
|
||||
}
|
||||
|
||||
return { command, env };
|
||||
}
|
||||
|
|
@ -87,11 +87,15 @@ describe("prepareSandboxClaudeProbeRuntime managed-config redaction", () => {
|
|||
const checkText = JSON.stringify(checks);
|
||||
expect(checkText).not.toContain(secret);
|
||||
expect(checkText).not.toContain("MANAGEDMARKER");
|
||||
// The diagnostic still reaches the server log, but the secret is redacted.
|
||||
// The diagnostic still reaches the server log, but it carries only the
|
||||
// fixed context and the allowlisted classification, never the raw error
|
||||
// text. The classification is `spawn_error`, because the materialization
|
||||
// step threw before the probe ran.
|
||||
expect(warnSpy).toHaveBeenCalledTimes(1);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(secret);
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
expect(loggedText).not.toContain("MANAGEDMARKER");
|
||||
expect(warnSpy.mock.calls[0]?.[1]).toMatchObject({ classification: "spawn_error" });
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,4 +1,7 @@
|
|||
import { afterEach, describe, expect, it, vi } from "vitest";
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { chmod, mkdtemp, rm, writeFile } from "node:fs/promises";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import type { AdapterExecutionTarget } from "@paperclipai/adapter-utils/execution-target";
|
||||
|
||||
const {
|
||||
|
|
@ -10,23 +13,30 @@ const {
|
|||
resolveAdapterExecutionTargetCwd,
|
||||
probeResult,
|
||||
} = vi.hoisted(() => {
|
||||
const probeResult: { value: { exitCode: number; stdout: string; stderr: string } } = {
|
||||
const probeResult: {
|
||||
value: { exitCode: number; stdout: string; stderr: string };
|
||||
throwError: Error | null;
|
||||
} = {
|
||||
value: { exitCode: 1, stdout: "", stderr: "" },
|
||||
throwError: null,
|
||||
};
|
||||
return {
|
||||
probeResult,
|
||||
ensureAdapterExecutionTargetDirectory: vi.fn(async () => {}),
|
||||
ensureAdapterExecutionTargetCommandResolvable: vi.fn(async () => {}),
|
||||
maybeRunSandboxInstallCommand: vi.fn(async () => null),
|
||||
runAdapterExecutionTargetProcess: vi.fn(async () => ({
|
||||
exitCode: probeResult.value.exitCode,
|
||||
signal: null,
|
||||
timedOut: false,
|
||||
stdout: probeResult.value.stdout,
|
||||
stderr: probeResult.value.stderr,
|
||||
pid: 123,
|
||||
startedAt: new Date().toISOString(),
|
||||
})),
|
||||
runAdapterExecutionTargetProcess: vi.fn(async () => {
|
||||
if (probeResult.throwError) throw probeResult.throwError;
|
||||
return {
|
||||
exitCode: probeResult.value.exitCode,
|
||||
signal: null,
|
||||
timedOut: false,
|
||||
stdout: probeResult.value.stdout,
|
||||
stderr: probeResult.value.stderr,
|
||||
pid: 123,
|
||||
startedAt: new Date().toISOString(),
|
||||
};
|
||||
}),
|
||||
describeAdapterExecutionTarget: vi.fn(() => "Daytona"),
|
||||
resolveAdapterExecutionTargetCwd: vi.fn(() => "/home/daytona/paperclip-workspace"),
|
||||
};
|
||||
|
|
@ -72,12 +82,13 @@ const initLine =
|
|||
|
||||
afterEach(() => {
|
||||
vi.clearAllMocks();
|
||||
probeResult.throwError = null;
|
||||
});
|
||||
|
||||
describe("claude sandbox hello probe diagnostics", () => {
|
||||
it("keeps the raw failure result out of every check and routes it to the log", async () => {
|
||||
it("keeps the raw failure result out of every check and out of the log", async () => {
|
||||
// The non-zero result event carries a marker. The check must not repeat the
|
||||
// marker, and the redacted diagnostic must reach the server log.
|
||||
// marker, and the log must carry only the allowlisted classification.
|
||||
const marker = "NONPATTERNMARKERfailure";
|
||||
probeResult.value = {
|
||||
exitCode: 1,
|
||||
|
|
@ -106,15 +117,17 @@ describe("claude sandbox hello probe diagnostics", () => {
|
|||
expect(checkText).not.toContain(marker);
|
||||
// The unhelpful init line must never reach a check either.
|
||||
expect(checkText).not.toContain('"subtype":"init"');
|
||||
// The raw diagnostic still reaches the server log.
|
||||
// The raw diagnostic never reaches the server log. The log carries only the
|
||||
// fixed context and an allowlisted classification.
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).toContain(marker);
|
||||
expect(loggedText).not.toContain(marker);
|
||||
expect(loggedText).toContain("nonzero_exit");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("keeps a stdout-fallback failure line out of every check", async () => {
|
||||
it("keeps a stdout-fallback failure line out of every check and out of the log", async () => {
|
||||
// The CLI dies before a result event, so the last non-init stdout line is
|
||||
// the diagnostic. The check must not repeat its marker.
|
||||
// the diagnostic. The check and the log must not repeat its marker.
|
||||
const marker = "NONPATTERNMARKERstdout";
|
||||
probeResult.value = {
|
||||
exitCode: 1,
|
||||
|
|
@ -138,13 +151,14 @@ describe("claude sandbox hello probe diagnostics", () => {
|
|||
expect(checkText).not.toContain(marker);
|
||||
expect(checkText).not.toContain('"subtype":"init"');
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).toContain(marker);
|
||||
expect(loggedText).not.toContain(marker);
|
||||
expect(loggedText).toContain("nonzero_exit");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("never copies a credential-bearing stderr failure line into a check", async () => {
|
||||
// A verbose CLI can print a credential to stderr on failure. The check must
|
||||
// not repeat it, and the server log must redact it.
|
||||
it("never copies a credential-bearing stderr failure line into a check or the log", async () => {
|
||||
// A verbose CLI can print a credential to stderr on failure. The check and
|
||||
// the log must not repeat it.
|
||||
const secret = "sk-ant-STDERRLEAK0123456789abcdef";
|
||||
probeResult.value = {
|
||||
exitCode: 1,
|
||||
|
|
@ -166,7 +180,8 @@ describe("claude sandbox hello probe diagnostics", () => {
|
|||
expect(checkText).not.toContain("STDERRLEAK");
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(secret);
|
||||
expect(loggedText).toContain("***REDACTED***");
|
||||
expect(loggedText).not.toContain("STDERRLEAK");
|
||||
expect(loggedText).toContain("nonzero_exit");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
|
|
@ -198,7 +213,8 @@ describe("claude sandbox hello probe diagnostics", () => {
|
|||
const checkText = JSON.stringify(result.checks);
|
||||
expect(checkText).not.toContain(marker);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).toContain(marker);
|
||||
expect(loggedText).not.toContain(marker);
|
||||
expect(loggedText).toContain("auth_required");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
|
|
@ -320,7 +336,82 @@ describe("claude sandbox hello probe diagnostics", () => {
|
|||
const checkText = JSON.stringify(result.checks);
|
||||
expect(checkText).not.toContain(marker);
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).toContain(marker);
|
||||
expect(loggedText).not.toContain(marker);
|
||||
expect(loggedText).toContain("unexpected_output");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("keeps an opaque credential marker and a proxy marker out of every check and the log", async () => {
|
||||
// The failure output carries two untrusted values that the pattern
|
||||
// sanitizer did not recognize: an opaque credential with no token shape and
|
||||
// a proxy URL. One rides in stdout, the other in stderr. Neither may reach a
|
||||
// check or the server log.
|
||||
const opaqueCredMarker = "OPAQUECREDMARKERnoshape";
|
||||
const proxyMarker = "http://user:pass@proxy.corp.internal:3128";
|
||||
probeResult.value = {
|
||||
exitCode: 7,
|
||||
stdout: [
|
||||
initLine,
|
||||
`{"type":"result","subtype":"error_during_execution","is_error":true,"result":"probe failed with ${opaqueCredMarker}","session_id":"abc"}`,
|
||||
].join("\n"),
|
||||
stderr: `proxy connect failed: ${proxyMarker}`,
|
||||
};
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
|
||||
const result = await testEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "cli", command: "claude" },
|
||||
executionTarget: sandboxTarget,
|
||||
environmentName: "Daytona",
|
||||
});
|
||||
|
||||
const checkText = JSON.stringify(result.checks);
|
||||
expect(checkText).not.toContain(opaqueCredMarker);
|
||||
expect(checkText).not.toContain("proxy.corp.internal");
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(opaqueCredMarker);
|
||||
expect(loggedText).not.toContain("proxy.corp.internal");
|
||||
// The log still carries the allowlisted classification and the safe exit
|
||||
// code, so the diagnostic stays useful.
|
||||
expect(loggedText).toContain("nonzero_exit");
|
||||
expect(warnSpy.mock.calls[0]?.[1]).toMatchObject({
|
||||
classification: "nonzero_exit",
|
||||
exitCode: 7,
|
||||
});
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
it("never copies a thrown CLI probe error into a check or the log", async () => {
|
||||
// A spawn or transport failure can throw an error whose text carries a
|
||||
// credential. Inject an opaque credential marker and a proxy marker through
|
||||
// the thrown error. The CLI lane has no catch around the hello probe call,
|
||||
// so the thrown error propagates to the caller, which owns it. No check and
|
||||
// no console.warn call inside this lane repeats either marker.
|
||||
const opaqueCredMarker = "OPAQUECREDMARKERnoshape";
|
||||
const proxyMarker = "http://user:pass@proxy.corp.internal:3128";
|
||||
probeResult.throwError = new Error(
|
||||
`probe spawn failed with ${opaqueCredMarker} via ${proxyMarker}`,
|
||||
);
|
||||
const warnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
|
||||
// The current contract propagates the thrown error. The lane builds no
|
||||
// Test-result check from the error, so the raw text cannot reach a check.
|
||||
await expect(
|
||||
testEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "cli", command: "claude" },
|
||||
executionTarget: sandboxTarget,
|
||||
environmentName: "Daytona",
|
||||
}),
|
||||
).rejects.toThrow(opaqueCredMarker);
|
||||
|
||||
// The lane never routes the raw error text to the server log. No
|
||||
// console.warn call repeats either marker.
|
||||
const loggedText = JSON.stringify(warnSpy.mock.calls);
|
||||
expect(loggedText).not.toContain(opaqueCredMarker);
|
||||
expect(loggedText).not.toContain("proxy.corp.internal");
|
||||
warnSpy.mockRestore();
|
||||
});
|
||||
|
||||
|
|
@ -445,3 +536,112 @@ describe("claude auth mode hints", () => {
|
|||
expect(result.checks.some((check) => check.code === "claude_oauth_token_configured")).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("claude CLI local hello probe hardening", () => {
|
||||
// Clear the host proxy and host auth variables so a local probe reads a
|
||||
// deterministic env regardless of the machine that runs the suite.
|
||||
const CLEARED_HOST_ENV_KEYS = [
|
||||
"HTTP_PROXY",
|
||||
"HTTPS_PROXY",
|
||||
"NO_PROXY",
|
||||
"http_proxy",
|
||||
"https_proxy",
|
||||
"no_proxy",
|
||||
"ANTHROPIC_API_KEY",
|
||||
"CLAUDE_CODE_OAUTH_TOKEN",
|
||||
"CLAUDE_CODE_USE_BEDROCK",
|
||||
"ANTHROPIC_BEDROCK_BASE_URL",
|
||||
];
|
||||
const successStdout = [
|
||||
initLine,
|
||||
'{"type":"result","subtype":"success","is_error":false,"result":"hello","session_id":"abc"}',
|
||||
].join("\n");
|
||||
|
||||
let tempDir: string | null = null;
|
||||
let claudePath = "";
|
||||
let savedPath: string | undefined;
|
||||
let savedEnv: Record<string, string | undefined> = {};
|
||||
|
||||
beforeEach(async () => {
|
||||
tempDir = await mkdtemp(path.join(os.tmpdir(), "paperclip-cli-localprobe-"));
|
||||
claudePath = path.join(tempDir, "claude");
|
||||
await writeFile(claudePath, "#!/bin/sh\nexit 0\n");
|
||||
await chmod(claudePath, 0o755);
|
||||
savedPath = process.env.PATH;
|
||||
process.env.PATH = tempDir;
|
||||
savedEnv = {};
|
||||
for (const key of CLEARED_HOST_ENV_KEYS) {
|
||||
savedEnv[key] = process.env[key];
|
||||
delete process.env[key];
|
||||
}
|
||||
// The mocked cwd resolver returns a sandbox path; the local probe reads it
|
||||
// as the cwd, so no host directory is touched.
|
||||
resolveAdapterExecutionTargetCwd.mockReturnValue("/home/daytona/paperclip-workspace");
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
process.env.PATH = savedPath;
|
||||
for (const [key, value] of Object.entries(savedEnv)) {
|
||||
if (value === undefined) delete process.env[key];
|
||||
else process.env[key] = value;
|
||||
}
|
||||
if (tempDir) await rm(tempDir, { recursive: true, force: true }).catch(() => undefined);
|
||||
tempDir = null;
|
||||
});
|
||||
|
||||
it("spawns the trusted resolved claude and drops hostile caller env for a local probe", async () => {
|
||||
probeResult.value = { exitCode: 0, stdout: successStdout, stderr: "" };
|
||||
process.env.HTTPS_PROXY = "http://trusted-proxy:8443";
|
||||
|
||||
await testEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: {
|
||||
engine: "cli",
|
||||
command: "/tmp/evil/claude",
|
||||
env: {
|
||||
ANTHROPIC_API_KEY: "keep-this-key",
|
||||
NODE_OPTIONS: "--require /hostile/evil.js",
|
||||
PATH: "/hostile/bin",
|
||||
LD_PRELOAD: "/hostile/evil.so",
|
||||
HTTP_PROXY: "http://caller-proxy:8080",
|
||||
},
|
||||
},
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
||||
const call = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as unknown[];
|
||||
const spawnedCommand = call[2] as string;
|
||||
const spawnedEnv = (call[4] as { env: Record<string, string> }).env;
|
||||
// The trusted resolved claude executable, never the caller command path.
|
||||
expect(spawnedCommand).toBe(claudePath);
|
||||
expect(spawnedCommand).not.toContain("/tmp/evil");
|
||||
// The approved key reaches the child; the hostile keys never do.
|
||||
expect(spawnedEnv.ANTHROPIC_API_KEY).toBe("keep-this-key");
|
||||
expect(spawnedEnv.NODE_OPTIONS).toBeUndefined();
|
||||
expect(spawnedEnv.PATH).toBeUndefined();
|
||||
expect(spawnedEnv.LD_PRELOAD).toBeUndefined();
|
||||
expect(spawnedEnv.HTTP_PROXY).toBeUndefined();
|
||||
// The trusted proxy reaches the child; the caller proxy never does.
|
||||
expect(spawnedEnv.HTTPS_PROXY).toBe("http://trusted-proxy:8443");
|
||||
expect(JSON.stringify(spawnedEnv)).not.toContain("caller-proxy");
|
||||
});
|
||||
|
||||
it("names the local host target on every result", async () => {
|
||||
probeResult.value = { exitCode: 0, stdout: successStdout, stderr: "" };
|
||||
|
||||
const result = await testEnvironment({
|
||||
companyId: "company-1",
|
||||
adapterType: "claude_local",
|
||||
config: { engine: "cli", command: "claude" },
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
});
|
||||
|
||||
const targetCheck = result.checks.find((check) => check.code === "claude_environment_target");
|
||||
expect(targetCheck).toBeTruthy();
|
||||
expect(targetCheck?.message).toContain("Paperclip host");
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -8,7 +8,6 @@ import {
|
|||
asBoolean,
|
||||
asNumber,
|
||||
asStringArray,
|
||||
parseJson,
|
||||
parseObject,
|
||||
ensurePathInEnv,
|
||||
} from "@paperclipai/adapter-utils/server-utils";
|
||||
|
|
@ -16,11 +15,9 @@ import {
|
|||
ensureAdapterExecutionTargetCommandResolvable,
|
||||
ensureAdapterExecutionTargetDirectory,
|
||||
runAdapterExecutionTargetProcess,
|
||||
describeAdapterExecutionTarget,
|
||||
resolveAdapterExecutionTargetCwd,
|
||||
} from "@paperclipai/adapter-utils/execution-target";
|
||||
import {
|
||||
describeClaudeFailure,
|
||||
detectClaudeLoginRequired,
|
||||
isClaudeProviderQuotaError,
|
||||
isClaudeTransientUpstreamError,
|
||||
|
|
@ -34,9 +31,11 @@ import { SANDBOX_INSTALL_COMMAND } from "../index.js";
|
|||
import { resolveClaudeExecutionEngineForRun, testClaudeAcpEnvironment } from "./acp.js";
|
||||
import { ADAPTER_AUTH_MISSING_CHECK_CODE } from "./auth-check.js";
|
||||
import {
|
||||
buildAdapterTestTargetCheck,
|
||||
buildClaudeLoginRequiredHint,
|
||||
logRedactedSandboxProbeDiagnostic,
|
||||
logSandboxProbeDiagnostic,
|
||||
} from "./probe-diagnostics.js";
|
||||
import { buildLocalAdapterTestProbeEnv } from "./probe-env.js";
|
||||
|
||||
function summarizeStatus(checks: AdapterEnvironmentCheck[]): AdapterEnvironmentTestResult["status"] {
|
||||
if (checks.some((check) => check.level === "error")) return "fail";
|
||||
|
|
@ -48,39 +47,6 @@ function isNonEmpty(value: unknown): value is string {
|
|||
return typeof value === "string" && value.trim().length > 0;
|
||||
}
|
||||
|
||||
function firstNonEmptyLine(text: string): string {
|
||||
return (
|
||||
text
|
||||
.split(/\r?\n/)
|
||||
.map((line) => line.trim())
|
||||
.find(Boolean) ?? ""
|
||||
);
|
||||
}
|
||||
|
||||
function lastNonInitStdoutLine(text: string): string {
|
||||
const lines = text
|
||||
.split(/\r?\n/)
|
||||
.map((line) => line.trim())
|
||||
.filter(Boolean);
|
||||
for (let index = lines.length - 1; index >= 0; index -= 1) {
|
||||
const line = lines[index]!;
|
||||
const parsed = parseJson(line);
|
||||
if (parsed && asString(parsed.type, "") === "system" && asString(parsed.subtype, "") === "init") {
|
||||
continue;
|
||||
}
|
||||
return line;
|
||||
}
|
||||
return "";
|
||||
}
|
||||
|
||||
function summarizeProbeDetail(stdout: string, stderr: string): string | null {
|
||||
const raw = firstNonEmptyLine(stderr) || lastNonInitStdoutLine(stdout);
|
||||
if (!raw) return null;
|
||||
const clean = raw.replace(/\s+/g, " ").trim();
|
||||
const max = 240;
|
||||
return clean.length > max ? `${clean.slice(0, max - 1)}…` : clean;
|
||||
}
|
||||
|
||||
export async function testEnvironment(
|
||||
ctx: AdapterEnvironmentTestContext,
|
||||
): Promise<AdapterEnvironmentTestResult> {
|
||||
|
|
@ -108,18 +74,13 @@ export async function testEnvironment(
|
|||
const targetIsRemote = target?.kind === "remote";
|
||||
const targetIsSandbox = target?.kind === "remote" && target.transport === "sandbox";
|
||||
const cwd = resolveAdapterExecutionTargetCwd(target, asString(config.cwd, ""), process.cwd());
|
||||
const targetLabel = targetIsRemote
|
||||
? ctx.environmentName ?? describeAdapterExecutionTarget(target)
|
||||
: null;
|
||||
const runId = `claude-envtest-${Date.now()}-${Math.random().toString(16).slice(2)}`;
|
||||
|
||||
if (targetLabel) {
|
||||
checks.push({
|
||||
code: "claude_environment_target",
|
||||
level: "info",
|
||||
message: `Probing inside environment: ${targetLabel}`,
|
||||
});
|
||||
}
|
||||
// Always name the target the Test probed, so a pass result never hides which
|
||||
// target it checked. A local probe reports the fixed host label.
|
||||
checks.push(
|
||||
buildAdapterTestTargetCheck({ targetIsRemote, environmentName: ctx.environmentName }),
|
||||
);
|
||||
|
||||
try {
|
||||
await ensureAdapterExecutionTargetDirectory(runId, target, cwd, {
|
||||
|
|
@ -146,6 +107,14 @@ export async function testEnvironment(
|
|||
for (const [key, value] of Object.entries(envConfig)) {
|
||||
if (typeof value === "string") env[key] = value;
|
||||
}
|
||||
// For a local probe, resolve the trusted `claude` executable and a
|
||||
// deny-by-default child env from the shared builder, so a hostile caller
|
||||
// value can neither select the executable nor reach the child. A remote
|
||||
// target keeps the caller command and env; the remote transport owns its own
|
||||
// env sanitization.
|
||||
const localProbe = targetIsRemote
|
||||
? null
|
||||
: await buildLocalAdapterTestProbeEnv({ callerEnv: env, trustedEnv: process.env });
|
||||
checks.push(
|
||||
...(await prepareSandboxClaudeProbeRuntime({
|
||||
runId,
|
||||
|
|
@ -254,6 +223,15 @@ export async function testEnvironment(
|
|||
detail: command,
|
||||
hint: "Use the `claude` CLI command to run the automatic login and installation probe.",
|
||||
});
|
||||
} else if (localProbe && !localProbe.command) {
|
||||
// The trusted server PATH holds no `claude`, so the local probe cannot
|
||||
// run. Report a warn, never a silent pass.
|
||||
checks.push({
|
||||
code: "claude_hello_probe_skipped_unresolved_command",
|
||||
level: "warn",
|
||||
message: "Skipped the Claude hello probe because `claude` is not installed on the Paperclip host.",
|
||||
hint: "Install the `claude` CLI on the Paperclip host, then retry the Test.",
|
||||
});
|
||||
} else {
|
||||
const model = asString(config.model, "").trim();
|
||||
const effort = asString(config.effort, "").trim();
|
||||
|
|
@ -312,14 +290,19 @@ export async function testEnvironment(
|
|||
asNumber(config.helloProbeTimeoutSec, targetIsSandbox ? 90 : 45),
|
||||
);
|
||||
|
||||
// A local probe uses the trusted resolved executable and the
|
||||
// deny-by-default child env. A remote probe uses the caller command and
|
||||
// env, because the remote transport owns its own env sanitization.
|
||||
const probeCommand = localProbe?.command ?? command;
|
||||
const probeEnv = localProbe ? localProbe.env : env;
|
||||
const probe = await runAdapterExecutionTargetProcess(
|
||||
runId,
|
||||
target,
|
||||
command,
|
||||
probeCommand,
|
||||
args,
|
||||
{
|
||||
cwd,
|
||||
env,
|
||||
env: probeEnv,
|
||||
timeoutSec: helloProbeTimeoutSec,
|
||||
graceSec: 5,
|
||||
stdin: "Respond with hello.",
|
||||
|
|
@ -343,11 +326,12 @@ export async function testEnvironment(
|
|||
hint: "Retry the probe. If this persists, verify Claude can run `Respond with hello` from this directory manually.",
|
||||
});
|
||||
} else if (loginMeta.requiresLogin) {
|
||||
// The raw probe output is untrusted. Route it to the log-only boundary
|
||||
// and return only a fixed public message and a safe hint.
|
||||
logRedactedSandboxProbeDiagnostic(
|
||||
// The raw probe output is untrusted. Log only the fixed context and the
|
||||
// allowlisted classification. Return only a fixed public message and a
|
||||
// safe hint.
|
||||
logSandboxProbeDiagnostic(
|
||||
"Claude CLI hello probe reported login required",
|
||||
summarizeProbeDetail(probe.stdout, probe.stderr),
|
||||
"auth_required",
|
||||
);
|
||||
checks.push({
|
||||
code: "claude_hello_probe_auth_required",
|
||||
|
|
@ -370,11 +354,12 @@ export async function testEnvironment(
|
|||
const summary = parsedStream.summary.trim();
|
||||
const hasHello = /\bhello\b/i.test(summary);
|
||||
if (!hasHello) {
|
||||
// The unexpected summary is untrusted probe output. Route it to the
|
||||
// log-only boundary and keep the check text fixed.
|
||||
logRedactedSandboxProbeDiagnostic(
|
||||
// The unexpected summary is untrusted probe output. Log only the fixed
|
||||
// context and the allowlisted classification. Keep the check text
|
||||
// fixed.
|
||||
logSandboxProbeDiagnostic(
|
||||
"Claude CLI hello probe returned unexpected output",
|
||||
summary,
|
||||
"unexpected_output",
|
||||
);
|
||||
}
|
||||
checks.push({
|
||||
|
|
@ -390,20 +375,12 @@ export async function testEnvironment(
|
|||
}),
|
||||
});
|
||||
} else {
|
||||
// Compose the richest raw diagnostic for the log. The real error lives
|
||||
// in the final `result` event (parsed) or, when the CLI dies before it
|
||||
// emits one, the last non-init stdout line — never the first line that
|
||||
// `summarizeProbeDetail` returns.
|
||||
const stdoutFallback = lastNonInitStdoutLine(probe.stdout);
|
||||
const failureDetail =
|
||||
(parsed ? describeClaudeFailure(parsed) : null) ||
|
||||
firstNonEmptyLine(probe.stderr) ||
|
||||
stdoutFallback ||
|
||||
summarizeProbeDetail(probe.stdout, probe.stderr) ||
|
||||
"";
|
||||
// The failure diagnostic is untrusted. Route it to the log-only
|
||||
// boundary and return only a fixed public message and hint.
|
||||
logRedactedSandboxProbeDiagnostic("Claude CLI hello probe failed", failureDetail);
|
||||
// The failure diagnostic is untrusted. Log only the fixed context, the
|
||||
// allowlisted classification, and the safe exit code. Return only a
|
||||
// fixed public message and hint.
|
||||
logSandboxProbeDiagnostic("Claude CLI hello probe failed", "nonzero_exit", {
|
||||
exitCode: probe.exitCode ?? null,
|
||||
});
|
||||
// Provider-quota exhaustion (usage/session limit) is classified
|
||||
// separately from generic transient upstream errors: auth works, the
|
||||
// subscription's usage window is just spent. Surface it as its own
|
||||
|
|
|
|||
|
|
@ -30,6 +30,8 @@ const mockSecretService = vi.hoisted(() => ({
|
|||
const mockEnvironmentService = vi.hoisted(() => ({
|
||||
getById: vi.fn(),
|
||||
releaseLease: vi.fn(),
|
||||
listBoundCompanyIds: vi.fn(async () => [] as string[]),
|
||||
findManagedSandboxEnvironment: vi.fn(async () => null as Record<string, unknown> | null),
|
||||
}));
|
||||
|
||||
const mockReleaseRunLease = vi.hoisted(() => vi.fn(async () => undefined));
|
||||
|
|
@ -44,6 +46,7 @@ const mockEnvironmentRuntime = vi.hoisted(() => ({
|
|||
const mockResolveEnvironmentExecutionTarget = vi.hoisted(() => vi.fn());
|
||||
const mockInstanceSettingsService = vi.hoisted(() => ({
|
||||
getGeneral: vi.fn(async () => ({ censorUsernameInLogs: false })),
|
||||
getExperimental: vi.fn(async () => ({ enableManagedSandboxOnly: false })),
|
||||
}));
|
||||
|
||||
vi.mock("../services/index.js", () => ({
|
||||
|
|
@ -140,6 +143,9 @@ describe("agent test-environment route", () => {
|
|||
driver: "sandbox",
|
||||
config: { provider: "fake-plugin" },
|
||||
});
|
||||
// Default to an instance-global environment with no company binding, so the
|
||||
// tenant-binding guard passes unless a test overrides it.
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue([]);
|
||||
mockEnvironmentRuntime.acquireRunLease.mockResolvedValue({
|
||||
lease: {
|
||||
id: "lease-1",
|
||||
|
|
@ -220,7 +226,10 @@ describe("agent test-environment route", () => {
|
|||
});
|
||||
|
||||
it("returns a diagnostic result instead of probing the host when the requested environment is missing", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValueOnce(null);
|
||||
// The route reads the environment more than once: the tenant-binding guard
|
||||
// loads it, then the execution-context resolver loads it. Return null for
|
||||
// every read so the missing-environment path is stable.
|
||||
mockEnvironmentService.getById.mockResolvedValue(null);
|
||||
const app = await createApp();
|
||||
|
||||
const res = await request(app)
|
||||
|
|
@ -282,6 +291,9 @@ describe("agent test-environment route", () => {
|
|||
expect(mockEnvironmentRuntime.acquireRunLease).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
applyCustomImageTemplate: true,
|
||||
// The Test lease re-checks the company binding, so a binding change
|
||||
// between the route guard and the lease cannot open a foreign sandbox.
|
||||
assertCompanyBinding: true,
|
||||
environment: expect.objectContaining({
|
||||
config: expect.objectContaining({
|
||||
reuseLease: false,
|
||||
|
|
@ -507,4 +519,172 @@ describe("agent test-environment route", () => {
|
|||
]);
|
||||
});
|
||||
});
|
||||
|
||||
describe("tenant-binding guard", () => {
|
||||
async function postForeignEnvironmentTest() {
|
||||
const app = await createApp();
|
||||
return request(app)
|
||||
.post("/api/companies/company-1/adapters/external_test/test-environment")
|
||||
.send({
|
||||
adapterConfig: { env: { FOO: "bar" } },
|
||||
environmentId: "11111111-1111-4111-8111-111111111111",
|
||||
});
|
||||
}
|
||||
|
||||
function expectCompanyMismatch(res: request.Response) {
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(403);
|
||||
expect(JSON.stringify(res.body)).toContain("environment_company_mismatch");
|
||||
// The guard rejects before any secret resolution, target resolution,
|
||||
// sandbox lease, or adapter test runs.
|
||||
expect(mockSecretService.normalizeAdapterConfigForPersistence).not.toHaveBeenCalled();
|
||||
expect(mockSecretService.resolveAdapterConfigForRuntime).not.toHaveBeenCalled();
|
||||
expect(mockEnvironmentRuntime.acquireRunLease).not.toHaveBeenCalled();
|
||||
expect(testEnvironmentSpy).not.toHaveBeenCalled();
|
||||
}
|
||||
|
||||
it("rejects an active environment bound to another company", async () => {
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue(["company-2"]);
|
||||
expectCompanyMismatch(await postForeignEnvironmentTest());
|
||||
});
|
||||
|
||||
it("rejects an archived environment bound to another company without revealing its status", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValue({
|
||||
id: "11111111-1111-4111-8111-111111111111",
|
||||
companyId: "company-2",
|
||||
name: "Sandbox QA",
|
||||
driver: "sandbox",
|
||||
status: "archived",
|
||||
config: { provider: "fake-plugin" },
|
||||
});
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue(["company-2"]);
|
||||
expectCompanyMismatch(await postForeignEnvironmentTest());
|
||||
});
|
||||
|
||||
it("rejects a disallowed-driver environment bound to another company without revealing its driver", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValue({
|
||||
id: "11111111-1111-4111-8111-111111111111",
|
||||
companyId: "company-2",
|
||||
name: "Plugin Env",
|
||||
driver: "plugin",
|
||||
status: "active",
|
||||
config: {},
|
||||
});
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue(["company-2"]);
|
||||
expectCompanyMismatch(await postForeignEnvironmentTest());
|
||||
});
|
||||
|
||||
it("allows an instance-global environment with no company binding", async () => {
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue([]);
|
||||
const res = await postForeignEnvironmentTest();
|
||||
// The guard passes and the route proceeds to secret resolution.
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(200);
|
||||
expect(mockSecretService.normalizeAdapterConfigForPersistence).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("allows an environment bound to the caller company", async () => {
|
||||
mockEnvironmentService.listBoundCompanyIds.mockResolvedValue(["company-1"]);
|
||||
const res = await postForeignEnvironmentTest();
|
||||
// The guard passes and the route proceeds to secret resolution.
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(200);
|
||||
expect(mockSecretService.normalizeAdapterConfigForPersistence).toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
describe("managed-sandbox-only redirect", () => {
|
||||
const localEnvironmentId = "33333333-3333-4333-8333-333333333333";
|
||||
const managedSandboxEnvironment = {
|
||||
id: "44444444-4444-4444-8444-444444444444",
|
||||
companyId: null,
|
||||
name: "Managed sandbox",
|
||||
driver: "sandbox",
|
||||
status: "active",
|
||||
config: { provider: "fake-plugin" },
|
||||
};
|
||||
const localEnvironment = {
|
||||
id: localEnvironmentId,
|
||||
companyId: null,
|
||||
name: "Local host",
|
||||
driver: "local",
|
||||
status: "active",
|
||||
config: {},
|
||||
};
|
||||
|
||||
it("redirects a local-environment Test onto the managed sandbox and never probes the host", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValue(localEnvironment);
|
||||
mockInstanceSettingsService.getExperimental.mockResolvedValue({
|
||||
enableManagedSandboxOnly: true,
|
||||
});
|
||||
mockEnvironmentService.findManagedSandboxEnvironment.mockResolvedValue(
|
||||
managedSandboxEnvironment,
|
||||
);
|
||||
mockResolveEnvironmentExecutionTarget.mockResolvedValueOnce({
|
||||
kind: "remote",
|
||||
transport: "sandbox",
|
||||
remoteCwd: "/home/user/paperclip-workspace",
|
||||
providerKey: "fake-plugin",
|
||||
runner: { execute: vi.fn() },
|
||||
});
|
||||
const app = await createApp();
|
||||
|
||||
const res = await request(app)
|
||||
.post("/api/companies/company-1/adapters/external_test/test-environment")
|
||||
.send({ adapterConfig: {}, environmentId: localEnvironmentId });
|
||||
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(200);
|
||||
// The Test leases and probes the managed sandbox the real run uses, not
|
||||
// the local host that the agent default still names.
|
||||
expect(mockEnvironmentRuntime.acquireRunLease).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
environment: expect.objectContaining({ id: managedSandboxEnvironment.id }),
|
||||
}),
|
||||
);
|
||||
expect(testEnvironmentSpy).toHaveBeenCalledTimes(1);
|
||||
expect(testEnvironmentSpy.mock.calls[0]?.[0]).toMatchObject({
|
||||
executionTarget: expect.objectContaining({ kind: "remote", transport: "sandbox" }),
|
||||
environmentName: "Managed sandbox",
|
||||
});
|
||||
});
|
||||
|
||||
it("fails closed when the policy is on and no managed sandbox environment exists", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValue(localEnvironment);
|
||||
mockInstanceSettingsService.getExperimental.mockResolvedValue({
|
||||
enableManagedSandboxOnly: true,
|
||||
});
|
||||
mockEnvironmentService.findManagedSandboxEnvironment.mockResolvedValue(null);
|
||||
const app = await createApp();
|
||||
|
||||
const res = await request(app)
|
||||
.post("/api/companies/company-1/adapters/external_test/test-environment")
|
||||
.send({ adapterConfig: {}, environmentId: localEnvironmentId });
|
||||
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(200);
|
||||
// No fall back to a host probe: the Test reports fail-closed.
|
||||
expect(testEnvironmentSpy).not.toHaveBeenCalled();
|
||||
expect(mockEnvironmentRuntime.acquireRunLease).not.toHaveBeenCalled();
|
||||
expect(res.body.status).toBe("fail");
|
||||
expect(res.body.checks).toEqual([
|
||||
expect.objectContaining({ code: "managed_sandbox_unavailable", level: "error" }),
|
||||
]);
|
||||
});
|
||||
|
||||
it("probes the local host when the managed-sandbox-only policy is off", async () => {
|
||||
mockEnvironmentService.getById.mockResolvedValue(localEnvironment);
|
||||
mockInstanceSettingsService.getExperimental.mockResolvedValue({
|
||||
enableManagedSandboxOnly: false,
|
||||
});
|
||||
const app = await createApp();
|
||||
|
||||
const res = await request(app)
|
||||
.post("/api/companies/company-1/adapters/external_test/test-environment")
|
||||
.send({ adapterConfig: {}, environmentId: localEnvironmentId });
|
||||
|
||||
expect(res.status, JSON.stringify(res.body)).toBe(200);
|
||||
// Legacy behavior: a local environment probes the host with no redirect
|
||||
// and no sandbox lease.
|
||||
expect(mockEnvironmentService.findManagedSandboxEnvironment).not.toHaveBeenCalled();
|
||||
expect(mockEnvironmentRuntime.acquireRunLease).not.toHaveBeenCalled();
|
||||
expect(testEnvironmentSpy).toHaveBeenCalledTimes(1);
|
||||
expect(testEnvironmentSpy.mock.calls[0]?.[0]?.executionTarget ?? null).toBeNull();
|
||||
});
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -36,6 +36,13 @@ const mockAccessService = vi.hoisted(() => ({
|
|||
const mockEnvironmentService = vi.hoisted(() => ({
|
||||
getById: vi.fn(),
|
||||
releaseLease: vi.fn(),
|
||||
// The tenant-binding guard reads the environment's bound company ids before it
|
||||
// reveals the driver or the status. An empty list marks an instance-global
|
||||
// environment, so the guard lets the same-company Test through.
|
||||
listBoundCompanyIds: vi.fn(async () => []),
|
||||
// The managed-sandbox-only redirect looks up the managed sandbox for a local
|
||||
// environment Test. These tests keep the policy off, so no managed row exists.
|
||||
findManagedSandboxEnvironment: vi.fn(async () => null),
|
||||
}));
|
||||
|
||||
const mockEnvironmentRuntime = vi.hoisted(() => ({
|
||||
|
|
@ -47,6 +54,7 @@ const mockEnvironmentRuntime = vi.hoisted(() => ({
|
|||
const mockResolveEnvironmentExecutionTarget = vi.hoisted(() => vi.fn(async () => null));
|
||||
const mockInstanceSettingsService = vi.hoisted(() => ({
|
||||
getGeneral: vi.fn(async () => ({ censorUsernameInLogs: false })),
|
||||
getExperimental: vi.fn(async () => ({ enableManagedSandboxOnly: false })),
|
||||
}));
|
||||
const mockRunClaudeLogin = vi.hoisted(() => vi.fn(async () => ({ ok: true })));
|
||||
|
||||
|
|
|
|||
|
|
@ -639,8 +639,8 @@ export function agentRoutes(
|
|||
};
|
||||
}
|
||||
|
||||
const environment = await environmentsSvc.getById(input.environmentId);
|
||||
if (!environment) {
|
||||
const requestedEnvironment = await environmentsSvc.getById(input.environmentId);
|
||||
if (!requestedEnvironment) {
|
||||
return {
|
||||
executionTarget: null,
|
||||
environmentName: null,
|
||||
|
|
@ -655,6 +655,40 @@ export function agentRoutes(
|
|||
};
|
||||
}
|
||||
|
||||
// Managed-sandbox-only policy: redirect a Test that would run on the local
|
||||
// host onto the platform-managed sandbox, the same as a real run does
|
||||
// (resolveExecutionWorkspaceEnvironmentId in heartbeat). Without this
|
||||
// redirect the Test probes the local host while the run executes in the
|
||||
// managed sandbox, so a passing Test validates the wrong execution target.
|
||||
// With no active managed sandbox the Test fails closed — never local.
|
||||
let environment = requestedEnvironment;
|
||||
if (requestedEnvironment.driver === "local") {
|
||||
const managedSandboxOnly =
|
||||
(await instanceSettings.getExperimental()).enableManagedSandboxOnly === true;
|
||||
if (managedSandboxOnly) {
|
||||
const managedSandboxEnvironment = await environmentsSvc.findManagedSandboxEnvironment(
|
||||
input.companyId,
|
||||
);
|
||||
if (!managedSandboxEnvironment) {
|
||||
return {
|
||||
executionTarget: null,
|
||||
environmentName: requestedEnvironment.name,
|
||||
fallbackChecks: [
|
||||
{
|
||||
code: "managed_sandbox_unavailable",
|
||||
level: "error",
|
||||
message:
|
||||
"This instance runs agents only in its platform-managed sandbox, but no active managed sandbox environment exists. The test did not run.",
|
||||
hint: "Restore the managed sandbox environment, then test again.",
|
||||
},
|
||||
],
|
||||
release: noopRelease,
|
||||
};
|
||||
}
|
||||
environment = managedSandboxEnvironment;
|
||||
}
|
||||
}
|
||||
|
||||
if (environment.driver === "local") {
|
||||
return {
|
||||
executionTarget: null,
|
||||
|
|
@ -748,6 +782,11 @@ export function agentRoutes(
|
|||
issueId: null,
|
||||
heartbeatRunId: null,
|
||||
persistedExecutionWorkspace: null,
|
||||
// Re-check the company binding atomically at lease time. The route
|
||||
// guard already rejected a foreign environment, but the binding could
|
||||
// change between the guard check and the lease acquire. This closes
|
||||
// that check-to-lease race so a foreign sandbox never gets a lease.
|
||||
assertCompanyBinding: true,
|
||||
// Apply the active custom-image template so the Test boots with the
|
||||
// operator's captured sandbox customizations and prepared image state,
|
||||
// matching what real agent runs use. Without this the test would
|
||||
|
|
@ -2346,6 +2385,39 @@ export function agentRoutes(
|
|||
res.json(detected);
|
||||
});
|
||||
|
||||
// The environment drivers the adapter Test route accepts. A local, SSH, or
|
||||
// sandbox environment can host a probe; a plugin environment cannot.
|
||||
const ADAPTER_TEST_ALLOWED_ENVIRONMENT_DRIVERS = ["local", "ssh", "sandbox"];
|
||||
|
||||
// The fail-closed tenant-binding guard for the adapter Test route. A caller
|
||||
// may name any instance environment by id, so the route must reject an
|
||||
// environment that binds to another company before it resolves secrets,
|
||||
// merges env, resolves the target, leases a sandbox, or runs the adapter
|
||||
// test. The guard checks the company binding BEFORE it validates the status
|
||||
// or the driver, so it never reveals the status or the driver of a foreign
|
||||
// environment. A same-company or an instance-global environment then gets the
|
||||
// shared driver and status validation.
|
||||
async function assertAdapterTestEnvironmentForCompany(
|
||||
companyId: string,
|
||||
environmentId: string,
|
||||
): Promise<void> {
|
||||
const environment = await environmentsSvc.getById(environmentId);
|
||||
if (!environment) {
|
||||
// A missing environment leaks no tenant state. The execution-context
|
||||
// resolver surfaces the existing environment_not_found check.
|
||||
return;
|
||||
}
|
||||
const boundCompanyIds = await environmentsSvc.listBoundCompanyIds(environmentId);
|
||||
if (boundCompanyIds.length > 0 && !boundCompanyIds.includes(companyId)) {
|
||||
throw forbidden("The selected environment belongs to another company.", {
|
||||
code: "environment_company_mismatch",
|
||||
});
|
||||
}
|
||||
await assertEnvironmentSelectionForCompany(environmentsSvc, companyId, environmentId, {
|
||||
allowedDrivers: ADAPTER_TEST_ALLOWED_ENVIRONMENT_DRIVERS,
|
||||
});
|
||||
}
|
||||
|
||||
router.post(
|
||||
"/companies/:companyId/adapters/:type/test-environment",
|
||||
validate(testAdapterEnvironmentSchema),
|
||||
|
|
@ -2362,6 +2434,11 @@ export function agentRoutes(
|
|||
typeof req.body?.environmentId === "string" && req.body.environmentId.trim().length > 0
|
||||
? (req.body.environmentId as string)
|
||||
: null;
|
||||
// Fail closed on a foreign environment before any secret resolution, env
|
||||
// merge, target resolution, sandbox lease, or adapter test runs.
|
||||
if (requestedEnvironmentId) {
|
||||
await assertAdapterTestEnvironmentForCompany(companyId, requestedEnvironmentId);
|
||||
}
|
||||
const normalizedAdapterConfig = await secretsSvc.normalizeAdapterConfigForPersistence(
|
||||
companyId,
|
||||
inputAdapterConfig,
|
||||
|
|
|
|||
|
|
@ -1263,6 +1263,73 @@ describe("AgentConfigForm environment selector", () => {
|
|||
expect(findButton(result.container, "Log in")).toBeFalsy();
|
||||
});
|
||||
|
||||
it("shows the Login button for an agent with no own environment under the managed-sandbox-only policy", async () => {
|
||||
// The agent has no own environment, so the login target resolves the same
|
||||
// way as the adapter Test target. The managed-sandbox-only policy redirects
|
||||
// that resolution from the hidden local environment to the managed sandbox.
|
||||
// The login affordance must read the managed sandbox, so it shows after the
|
||||
// auth-missing check. A login target that stayed local would hide the panel
|
||||
// for the target the real run uses.
|
||||
mockInstanceSettingsApi.getExperimental.mockResolvedValue({
|
||||
enableEnvironments: true,
|
||||
enableManagedSandboxOnly: true,
|
||||
});
|
||||
mockAgentsApi.testEnvironment.mockResolvedValue(CLAUDE_AUTH_MISSING_RESULT);
|
||||
const result = await renderForm(
|
||||
[
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
name: "Local",
|
||||
driver: "local",
|
||||
metadata: { defaultForInstance: true },
|
||||
}),
|
||||
makeEnvironment({
|
||||
id: "managed-1",
|
||||
name: "Managed",
|
||||
driver: "sandbox",
|
||||
config: { provider: "daytona" },
|
||||
metadata: { managedByPaperclip: true },
|
||||
}),
|
||||
],
|
||||
{ adapterType: "claude_local", defaultEnvironmentId: null },
|
||||
{ showAdapterTestEnvironmentButton: true },
|
||||
);
|
||||
roots.push(result.root);
|
||||
|
||||
expect(findButton(result.container, "Log in")).toBeFalsy();
|
||||
|
||||
await runTest(result.container);
|
||||
|
||||
expect(findButton(result.container, "Log in")).toBeTruthy();
|
||||
});
|
||||
|
||||
it("keeps the Login button hidden under the managed-sandbox-only policy when no managed sandbox is available", async () => {
|
||||
// The policy is on, but no managed sandbox environment exists, so the login
|
||||
// target resolution fails closed. The render catches that failure and
|
||||
// resolves no login environment, so the affordance stays hidden. The Test
|
||||
// surfaces the same case as a fail-closed error.
|
||||
mockInstanceSettingsApi.getExperimental.mockResolvedValue({
|
||||
enableEnvironments: true,
|
||||
enableManagedSandboxOnly: true,
|
||||
});
|
||||
mockAgentsApi.testEnvironment.mockResolvedValue(CLAUDE_AUTH_MISSING_RESULT);
|
||||
const result = await renderForm(
|
||||
[
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
name: "Local",
|
||||
driver: "local",
|
||||
metadata: { defaultForInstance: true },
|
||||
}),
|
||||
],
|
||||
{ adapterType: "claude_local", defaultEnvironmentId: null },
|
||||
{ showAdapterTestEnvironmentButton: true },
|
||||
);
|
||||
roots.push(result.root);
|
||||
|
||||
expect(findButton(result.container, "Log in")).toBeFalsy();
|
||||
});
|
||||
|
||||
it("starts a login session for the effective sandbox and shows the code and the authentication URL", async () => {
|
||||
mockAgentsApi.testEnvironment.mockResolvedValue(AUTH_MISSING_RESULT);
|
||||
const result = await renderCodexSandbox();
|
||||
|
|
|
|||
|
|
@ -32,7 +32,11 @@ import { Button } from "@/components/ui/button";
|
|||
import { FolderOpen, Heart, ChevronDown, X, Copy, Check, ExternalLink, Loader2, TriangleAlert } from "lucide-react";
|
||||
import { asBoolean, asFiniteNumber, asObject, cn } from "../lib/utils";
|
||||
import { copyTextToClipboard } from "../lib/clipboard";
|
||||
import { resolveAdapterTestEnvironmentId } from "../lib/adapter-test-environment";
|
||||
import {
|
||||
resolveAdapterTestEnvironmentId,
|
||||
resolveLocalDefaultEnvironmentId,
|
||||
resolveManagedSandboxEnvironmentId,
|
||||
} from "../lib/adapter-test-environment";
|
||||
import { extractModelName, extractProviderId } from "../lib/model-utils";
|
||||
import { queryKeys } from "../lib/queryKeys";
|
||||
import { useCompany } from "../context/CompanyContext";
|
||||
|
|
@ -581,17 +585,43 @@ export function AgentConfigForm(props: AgentConfigFormProps) {
|
|||
);
|
||||
|
||||
// The environment a login session runs in. It mirrors the Test resolution: the
|
||||
// agent's own environment wins, otherwise the instance default. The login
|
||||
// affordance shows only when this environment is a sandbox, because the
|
||||
// canonical auth-missing check comes only from a sandbox target.
|
||||
const effectiveLoginEnvironmentId = useMemo(
|
||||
() =>
|
||||
resolveAdapterTestEnvironmentId({
|
||||
// agent's own environment wins, otherwise the instance default, otherwise the
|
||||
// local default. The login affordance shows only when this environment is a
|
||||
// sandbox, because the canonical auth-missing check comes only from a sandbox
|
||||
// target.
|
||||
//
|
||||
// The resolution passes the same managed-sandbox-only policy inputs as the
|
||||
// adapter Test target, so both resolve to the same environment. Under the
|
||||
// policy a resolution that lands on the local environment redirects to the
|
||||
// managed sandbox the real run uses. Without the redirect the login target
|
||||
// stays local while the Test and the real run use the managed sandbox, so the
|
||||
// login affordance reads the wrong target. The resolver throws when the policy
|
||||
// is on but no managed sandbox is available; a render must not throw, so this
|
||||
// resolution catches that case and resolves no login environment. The Test
|
||||
// mutation surfaces the same case as a fail-closed error.
|
||||
const effectiveLoginEnvironmentId = useMemo(() => {
|
||||
try {
|
||||
return resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: rawCurrentDefaultEnvironmentId || null,
|
||||
instanceDefaultEnvironmentId: instanceSettings?.defaultEnvironmentId ?? null,
|
||||
}),
|
||||
[rawCurrentDefaultEnvironmentId, instanceSettings?.defaultEnvironmentId],
|
||||
);
|
||||
localDefaultEnvironmentId: resolveLocalDefaultEnvironmentId(environments),
|
||||
managedSandboxOnly: experimentalSettings?.enableManagedSandboxOnly === true,
|
||||
managedSandboxEnvironmentId: resolveManagedSandboxEnvironmentId(environments),
|
||||
// The policy hides the local environment, so an agent default that still
|
||||
// points at the hidden local row names no visible environment. Pass the
|
||||
// visible ids so the resolver redirects that stale local default to the
|
||||
// managed sandbox instead of the hidden local id.
|
||||
visibleEnvironmentIds: environments.map((environment) => environment.id),
|
||||
});
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}, [
|
||||
rawCurrentDefaultEnvironmentId,
|
||||
instanceSettings?.defaultEnvironmentId,
|
||||
environments,
|
||||
experimentalSettings?.enableManagedSandboxOnly,
|
||||
]);
|
||||
const effectiveLoginEnvironment = useMemo(
|
||||
() => environments.find((environment) => environment.id === effectiveLoginEnvironmentId) ?? null,
|
||||
[environments, effectiveLoginEnvironmentId],
|
||||
|
|
@ -858,21 +888,57 @@ export function AgentConfigForm(props: AgentConfigFormProps) {
|
|||
// the exact false command-not-found failure this resolution exists to
|
||||
// fix. Agents with their own environment never need the settings.
|
||||
let settings = instanceSettings;
|
||||
if (!rawCurrentDefaultEnvironmentId && settings === undefined) {
|
||||
let environmentList = environments;
|
||||
let managedSandboxOnly = experimentalSettings?.enableManagedSandboxOnly === true;
|
||||
if (!rawCurrentDefaultEnvironmentId) {
|
||||
// The agent has no own environment, so the Test resolves the instance
|
||||
// default, the local default, or the managed sandbox. Resolve the
|
||||
// settings, the environment list, and the managed-sandbox-only policy
|
||||
// here, because the render-time queries can be unsettled, or the
|
||||
// environments query can be disabled under the managed-sandbox-only
|
||||
// policy. A failure surfaces an honest error, not a silent host probe
|
||||
// that reports a false result.
|
||||
try {
|
||||
settings = await queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.instance.settings,
|
||||
queryFn: () => instanceSettingsApi.get(),
|
||||
});
|
||||
const [resolvedSettings, resolvedEnvironments, resolvedExperimental] =
|
||||
await Promise.all([
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.instance.settings,
|
||||
queryFn: () => instanceSettingsApi.get(),
|
||||
}),
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.environments.list(selectedCompanyId),
|
||||
queryFn: () => environmentsApi.list(selectedCompanyId),
|
||||
}),
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.instance.experimentalSettings,
|
||||
queryFn: () => instanceSettingsApi.getExperimental(),
|
||||
}),
|
||||
]);
|
||||
settings = resolvedSettings;
|
||||
environmentList = resolvedEnvironments;
|
||||
managedSandboxOnly = resolvedExperimental?.enableManagedSandboxOnly === true;
|
||||
} catch {
|
||||
throw new Error(
|
||||
"Could not load instance settings to determine which environment to test in. Retry the test.",
|
||||
"Could not load environment settings to determine which environment to test in. Retry the test.",
|
||||
);
|
||||
}
|
||||
}
|
||||
// Mirror the server run-time resolution, including the managed-sandbox-only
|
||||
// redirect: when the resolution lands on the local environment and the
|
||||
// policy is on, probe the managed sandbox the real run uses instead. The
|
||||
// resolver throws when no managed sandbox is available, which the mutation
|
||||
// surfaces as a fail-closed error rather than a local host probe.
|
||||
const environmentId = resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: rawCurrentDefaultEnvironmentId || null,
|
||||
instanceDefaultEnvironmentId: settings?.defaultEnvironmentId ?? null,
|
||||
localDefaultEnvironmentId: resolveLocalDefaultEnvironmentId(environmentList),
|
||||
managedSandboxOnly,
|
||||
managedSandboxEnvironmentId: resolveManagedSandboxEnvironmentId(environmentList),
|
||||
// The policy hides the local environment, so an agent default that still
|
||||
// points at the hidden local row names no visible environment. Pass the
|
||||
// visible ids so the resolver redirects that stale local default to the
|
||||
// managed sandbox instead of sending the hidden local id to the server.
|
||||
visibleEnvironmentIds: environmentList.map((environment) => environment.id),
|
||||
});
|
||||
const testResults: Array<{ label: string; model: string | null; result: AdapterEnvironmentTestResult }> = [
|
||||
{
|
||||
|
|
|
|||
|
|
@ -41,6 +41,15 @@ const mockAgentsApi = vi.hoisted(() => ({
|
|||
testEnvironment: vi.fn(),
|
||||
}));
|
||||
const mockCompaniesApi = vi.hoisted(() => ({ create: vi.fn() }));
|
||||
// The hire path resolves the Test environment before it probes: it reads the
|
||||
// environment list, the instance settings, and the experimental settings. The
|
||||
// test stubs these so the resolution settles on the local default, the same as
|
||||
// a real run with no instance default.
|
||||
const mockEnvironmentsApi = vi.hoisted(() => ({ list: vi.fn() }));
|
||||
const mockInstanceSettingsApi = vi.hoisted(() => ({
|
||||
get: vi.fn(),
|
||||
getExperimental: vi.fn(),
|
||||
}));
|
||||
|
||||
const routerState = vi.hoisted(() => ({ pathname: "/" }));
|
||||
const dialogState = vi.hoisted(() => ({
|
||||
|
|
@ -66,6 +75,8 @@ vi.mock("../api/agents", () => ({ agentsApi: mockAgentsApi }));
|
|||
vi.mock("../api/approvals", () => ({ approvalsApi: { create: vi.fn() } }));
|
||||
vi.mock("../api/issues", () => ({ issuesApi: { create: vi.fn() } }));
|
||||
vi.mock("../api/projects", () => ({ projectsApi: { list: vi.fn(), create: vi.fn() } }));
|
||||
vi.mock("../api/environments", () => ({ environmentsApi: mockEnvironmentsApi }));
|
||||
vi.mock("../api/instanceSettings", () => ({ instanceSettingsApi: mockInstanceSettingsApi }));
|
||||
|
||||
vi.mock("@/lib/router", () => ({
|
||||
useLocation: () => ({ pathname: routerState.pathname }),
|
||||
|
|
@ -192,6 +203,11 @@ describe("OnboardingWizard — which step it lands on", () => {
|
|||
checks: [],
|
||||
testedAt: new Date("2026-03-02T00:00:00Z").toISOString(),
|
||||
});
|
||||
mockEnvironmentsApi.list.mockResolvedValue([]);
|
||||
mockInstanceSettingsApi.get.mockResolvedValue({ defaultEnvironmentId: null });
|
||||
mockInstanceSettingsApi.getExperimental.mockResolvedValue({
|
||||
enableManagedSandboxOnly: false,
|
||||
});
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
|
|
|
|||
|
|
@ -1,6 +1,11 @@
|
|||
import { useEffect, useState, useMemo, useRef } from "react";
|
||||
import type { CSSProperties } from "react";
|
||||
import { useQuery, useQueryClient } from "@tanstack/react-query";
|
||||
import type { AdapterEnvironmentTestResult } from "@paperclipai/shared";
|
||||
import type {
|
||||
AdapterEnvironmentTestResult,
|
||||
Environment,
|
||||
InstanceSettings,
|
||||
} from "@paperclipai/shared";
|
||||
import { useLocation, useNavigate, useParams } from "@/lib/router";
|
||||
import { useDialog } from "../context/DialogContext";
|
||||
import { useCompany } from "../context/CompanyContext";
|
||||
|
|
@ -11,6 +16,13 @@ import { agentsApi } from "../api/agents";
|
|||
import { approvalsApi } from "../api/approvals";
|
||||
import { issuesApi } from "../api/issues";
|
||||
import { projectsApi } from "../api/projects";
|
||||
import { environmentsApi } from "../api/environments";
|
||||
import { instanceSettingsApi } from "../api/instanceSettings";
|
||||
import {
|
||||
resolveAdapterTestEnvironmentId,
|
||||
resolveLocalDefaultEnvironmentId,
|
||||
resolveManagedSandboxEnvironmentId,
|
||||
} from "../lib/adapter-test-environment";
|
||||
import { queryKeys } from "../lib/queryKeys";
|
||||
import { Dialog, DialogPortal } from "@/components/ui/dialog";
|
||||
import {
|
||||
|
|
@ -974,11 +986,61 @@ function OnboardingWizardInner({
|
|||
setAdapterEnvLoading(true);
|
||||
setAdapterEnvError(null);
|
||||
try {
|
||||
// Probe the environment a real run would use, so the Test matches a real
|
||||
// run. The wizard has no agent yet, so the agent-default tier is always
|
||||
// null; resolve the instance default and the instance local default. A
|
||||
// settings-resolution failure surfaces an error instead of a silent host
|
||||
// probe, which would report a false result.
|
||||
let environmentList: Environment[];
|
||||
let settings: InstanceSettings;
|
||||
let managedSandboxOnly: boolean;
|
||||
try {
|
||||
const [list, generalSettings, experimentalSettings] = await Promise.all([
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.environments.list(createdCompanyId),
|
||||
queryFn: () => environmentsApi.list(createdCompanyId),
|
||||
}),
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.instance.settings,
|
||||
queryFn: () => instanceSettingsApi.get(),
|
||||
}),
|
||||
queryClient.ensureQueryData({
|
||||
queryKey: queryKeys.instance.experimentalSettings,
|
||||
queryFn: () => instanceSettingsApi.getExperimental(),
|
||||
}),
|
||||
]);
|
||||
environmentList = list;
|
||||
settings = generalSettings;
|
||||
managedSandboxOnly = experimentalSettings?.enableManagedSandboxOnly === true;
|
||||
} catch {
|
||||
setAdapterEnvError(
|
||||
"Could not load environment settings to determine which environment to test in. Retry the test.",
|
||||
);
|
||||
return null;
|
||||
}
|
||||
// Mirror the server run-time resolution, including the managed-sandbox-only
|
||||
// redirect: when the resolution lands on the local environment and the
|
||||
// policy is on, probe the managed sandbox the real run uses instead. The
|
||||
// resolver throws when no managed sandbox is available, which the outer
|
||||
// catch surfaces as a fail-closed error rather than a local host probe.
|
||||
const environmentId = resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: settings?.defaultEnvironmentId ?? null,
|
||||
localDefaultEnvironmentId: resolveLocalDefaultEnvironmentId(environmentList),
|
||||
managedSandboxOnly,
|
||||
managedSandboxEnvironmentId: resolveManagedSandboxEnvironmentId(environmentList),
|
||||
// The policy hides the local environment, so an instance default that
|
||||
// still points at the hidden local row names no visible environment.
|
||||
// Pass the visible ids so the resolver redirects that stale local
|
||||
// default to the managed sandbox instead of sending the hidden local id.
|
||||
visibleEnvironmentIds: environmentList.map((environment) => environment.id),
|
||||
});
|
||||
const result = await agentsApi.testEnvironment(
|
||||
createdCompanyId,
|
||||
adapterType,
|
||||
{
|
||||
adapterConfig: adapterConfigOverride ?? buildAdapterConfig()
|
||||
adapterConfig: adapterConfigOverride ?? buildAdapterConfig(),
|
||||
environmentId,
|
||||
}
|
||||
);
|
||||
setAdapterEnvResult(result);
|
||||
|
|
@ -1160,6 +1222,14 @@ function OnboardingWizardInner({
|
|||
if (isLocalAdapter) {
|
||||
const result = adapterEnvResult ?? (await runAdapterEnvironmentTest());
|
||||
if (!result) return;
|
||||
// Block the hire on a failed environment test. A pass or a warn may
|
||||
// proceed; a fail means the agent cannot run as configured.
|
||||
if (result.status === "fail") {
|
||||
setError(
|
||||
"The environment test failed. Fix the reported checks before you hire this agent.",
|
||||
);
|
||||
return;
|
||||
}
|
||||
}
|
||||
|
||||
const hire = await agentsApi.hire(createdCompanyId, {
|
||||
|
|
@ -2030,9 +2100,21 @@ function OnboardingWizardInner({
|
|||
|
||||
{adapterEnvResult &&
|
||||
adapterEnvResult.status === "pass" ? (
|
||||
<div className="flex items-center gap-2 rounded-md border border-green-300 dark:border-green-500/40 bg-green-50 dark:bg-green-500/10 px-3 py-2 text-xs text-green-700 dark:text-green-300 animate-in fade-in slide-in-from-bottom-1 duration-300">
|
||||
<Check className="h-3.5 w-3.5 shrink-0" />
|
||||
<span className="font-medium">Passed</span>
|
||||
<div className="space-y-2 animate-in fade-in slide-in-from-bottom-1 duration-300">
|
||||
{/* Use the shared status-chip helper with the done
|
||||
status hue, so the pass banner derives its fill,
|
||||
text, and border from the design tokens in both
|
||||
modes instead of raw color values. */}
|
||||
<div
|
||||
className="status-chip flex items-center gap-2 rounded-md border px-2.5 py-2 text-(length:--text-micro)"
|
||||
style={{ "--sc": "var(--status-task-done)" } as CSSProperties}
|
||||
>
|
||||
<Check className="size-3.5 shrink-0" />
|
||||
<span className="font-medium">Passed</span>
|
||||
</div>
|
||||
{/* Show the checks on a pass too, so the target and the
|
||||
auth signals stay visible before the hire. */}
|
||||
<AdapterEnvironmentResult result={adapterEnvResult} />
|
||||
</div>
|
||||
) : adapterEnvResult ? (
|
||||
<AdapterEnvironmentResult result={adapterEnvResult} />
|
||||
|
|
|
|||
|
|
@ -1,6 +1,28 @@
|
|||
import { describe, expect, it } from "vitest";
|
||||
import type { Environment } from "@paperclipai/shared";
|
||||
|
||||
import { resolveAdapterTestEnvironmentId } from "./adapter-test-environment";
|
||||
import {
|
||||
ManagedSandboxUnavailableForTestError,
|
||||
resolveAdapterTestEnvironmentId,
|
||||
resolveLocalDefaultEnvironmentId,
|
||||
resolveManagedSandboxEnvironmentId,
|
||||
} from "./adapter-test-environment";
|
||||
|
||||
function makeEnvironment(overrides: Partial<Environment>): Environment {
|
||||
return {
|
||||
id: "env-id",
|
||||
name: "Env",
|
||||
description: null,
|
||||
driver: "sandbox",
|
||||
status: "active",
|
||||
config: {},
|
||||
envVars: {},
|
||||
metadata: null,
|
||||
createdAt: new Date(0),
|
||||
updatedAt: new Date(0),
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
describe("resolveAdapterTestEnvironmentId", () => {
|
||||
it("prefers the agent's own environment", () => {
|
||||
|
|
@ -8,6 +30,7 @@ describe("resolveAdapterTestEnvironmentId", () => {
|
|||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "agent-env",
|
||||
instanceDefaultEnvironmentId: "instance-env",
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
}),
|
||||
).toBe("agent-env");
|
||||
});
|
||||
|
|
@ -21,28 +44,261 @@ describe("resolveAdapterTestEnvironmentId", () => {
|
|||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: "instance-env",
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
}),
|
||||
).toBe("instance-env");
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "",
|
||||
instanceDefaultEnvironmentId: "instance-env",
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
}),
|
||||
).toBe("instance-env");
|
||||
});
|
||||
|
||||
it("returns null (host probe) when neither is set", () => {
|
||||
it("falls back to the local default when neither an agent nor an instance default is set", () => {
|
||||
// The server resolves a run with no agent or instance default to the local
|
||||
// default environment. The Test must probe the same environment, not the
|
||||
// host, so a Test result matches a real run.
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
}),
|
||||
).toBe("local-env");
|
||||
});
|
||||
|
||||
it("returns null (host probe) when no tier is set", () => {
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: undefined,
|
||||
instanceDefaultEnvironmentId: undefined,
|
||||
localDefaultEnvironmentId: undefined,
|
||||
}),
|
||||
).toBeNull();
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "",
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: null,
|
||||
}),
|
||||
).toBeNull();
|
||||
});
|
||||
|
||||
it("redirects a local-default resolution to the managed sandbox under managed-sandbox-only", () => {
|
||||
// The real run redirects a local resolution to the managed sandbox, so the
|
||||
// Test must probe the managed sandbox, not the local default.
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
}),
|
||||
).toBe("managed-env");
|
||||
});
|
||||
|
||||
it("redirects a no-tier resolution to the managed sandbox under managed-sandbox-only", () => {
|
||||
// The policy hides the local tier, so the list holds no local row and the
|
||||
// resolution lands on no tier. The real run still resolves to local and
|
||||
// redirects to the managed sandbox, so the Test does the same.
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: null,
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
}),
|
||||
).toBe("managed-env");
|
||||
});
|
||||
|
||||
it("fails closed when managed-sandbox-only is on but no managed sandbox exists", () => {
|
||||
// The real run fails closed, so the Test must not fall back to a local host
|
||||
// probe that reports a result for a target the run never uses.
|
||||
expect(() =>
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: null,
|
||||
}),
|
||||
).toThrow(ManagedSandboxUnavailableForTestError);
|
||||
});
|
||||
|
||||
it("leaves a non-local default untouched under managed-sandbox-only", () => {
|
||||
// The policy hides local; it does not forbid an ssh or a user sandbox that
|
||||
// an agent or instance default names, so the Test probes that environment.
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: "ssh-env",
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
}),
|
||||
).toBe("ssh-env");
|
||||
});
|
||||
|
||||
it("ignores the managed-sandbox redirect when the policy is off", () => {
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: null,
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: "local-env",
|
||||
managedSandboxOnly: false,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
}),
|
||||
).toBe("local-env");
|
||||
});
|
||||
|
||||
it("redirects an agent default that names the hidden local row to the managed sandbox", () => {
|
||||
// The policy hides the local environment, so the client list holds no local
|
||||
// row and the local-default lookup is null. An agent default that still
|
||||
// names the hidden local row names no visible environment, so the resolver
|
||||
// treats it as a local resolution and redirects to the managed sandbox.
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "hidden-local-env",
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: null,
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
visibleEnvironmentIds: ["managed-env", "ssh-env"],
|
||||
}),
|
||||
).toBe("managed-env");
|
||||
});
|
||||
|
||||
it("fails closed when an agent default names the hidden local row and no managed sandbox exists", () => {
|
||||
expect(() =>
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "hidden-local-env",
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: null,
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: null,
|
||||
visibleEnvironmentIds: ["ssh-env"],
|
||||
}),
|
||||
).toThrow(ManagedSandboxUnavailableForTestError);
|
||||
});
|
||||
|
||||
it("leaves a visible non-local default untouched when a visibility list is supplied", () => {
|
||||
expect(
|
||||
resolveAdapterTestEnvironmentId({
|
||||
agentDefaultEnvironmentId: "ssh-env",
|
||||
instanceDefaultEnvironmentId: null,
|
||||
localDefaultEnvironmentId: null,
|
||||
managedSandboxOnly: true,
|
||||
managedSandboxEnvironmentId: "managed-env",
|
||||
visibleEnvironmentIds: ["ssh-env", "managed-env"],
|
||||
}),
|
||||
).toBe("ssh-env");
|
||||
});
|
||||
});
|
||||
|
||||
describe("resolveManagedSandboxEnvironmentId", () => {
|
||||
it("finds the non-local platform-managed environment", () => {
|
||||
const environments = [
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
driver: "local",
|
||||
metadata: { managedByPaperclip: true, defaultForInstance: true },
|
||||
}),
|
||||
makeEnvironment({
|
||||
id: "sandbox-1",
|
||||
driver: "sandbox",
|
||||
metadata: { managedByPaperclip: true },
|
||||
}),
|
||||
];
|
||||
// The local-default row also carries the managed stamp, so the resolver must
|
||||
// exclude the local driver and return the sandbox row.
|
||||
expect(resolveManagedSandboxEnvironmentId(environments)).toBe("sandbox-1");
|
||||
});
|
||||
|
||||
it("returns null when no managed sandbox environment exists", () => {
|
||||
const environments = [
|
||||
makeEnvironment({ id: "sandbox-1", driver: "sandbox", metadata: null }),
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
driver: "local",
|
||||
metadata: { managedByPaperclip: true, defaultForInstance: true },
|
||||
}),
|
||||
];
|
||||
expect(resolveManagedSandboxEnvironmentId(environments)).toBeNull();
|
||||
expect(resolveManagedSandboxEnvironmentId([])).toBeNull();
|
||||
expect(resolveManagedSandboxEnvironmentId(null)).toBeNull();
|
||||
});
|
||||
|
||||
it("skips an archived managed sandbox", () => {
|
||||
// The server run-time resolver requires an active managed sandbox, so the
|
||||
// Test resolution must not select an archived one. An archived sandbox row
|
||||
// still carries the managed stamp, so the status guard is the only filter
|
||||
// that excludes it. A real run rejects an archived sandbox, so the Test must
|
||||
// reject it too.
|
||||
const environments = [
|
||||
makeEnvironment({
|
||||
id: "sandbox-archived",
|
||||
driver: "sandbox",
|
||||
status: "archived",
|
||||
metadata: { managedByPaperclip: true },
|
||||
}),
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
driver: "local",
|
||||
metadata: { managedByPaperclip: true, defaultForInstance: true },
|
||||
}),
|
||||
];
|
||||
expect(resolveManagedSandboxEnvironmentId(environments)).toBeNull();
|
||||
});
|
||||
|
||||
it("selects the active managed sandbox and skips an archived one", () => {
|
||||
// When both an archived and an active managed sandbox exist, the resolver
|
||||
// returns the active row, matching the server run-time resolver.
|
||||
const environments = [
|
||||
makeEnvironment({
|
||||
id: "sandbox-archived",
|
||||
driver: "sandbox",
|
||||
status: "archived",
|
||||
metadata: { managedByPaperclip: true },
|
||||
}),
|
||||
makeEnvironment({
|
||||
id: "sandbox-active",
|
||||
driver: "sandbox",
|
||||
status: "active",
|
||||
metadata: { managedByPaperclip: true },
|
||||
}),
|
||||
];
|
||||
expect(resolveManagedSandboxEnvironmentId(environments)).toBe("sandbox-active");
|
||||
});
|
||||
});
|
||||
|
||||
describe("resolveLocalDefaultEnvironmentId", () => {
|
||||
it("finds the local-driver instance-default environment", () => {
|
||||
const environments = [
|
||||
makeEnvironment({ id: "sandbox-1", driver: "sandbox" }),
|
||||
makeEnvironment({
|
||||
id: "local-1",
|
||||
driver: "local",
|
||||
metadata: { managedByPaperclip: true, defaultForInstance: true },
|
||||
}),
|
||||
];
|
||||
expect(resolveLocalDefaultEnvironmentId(environments)).toBe("local-1");
|
||||
});
|
||||
|
||||
it("ignores a local environment that is not the instance default", () => {
|
||||
const environments = [
|
||||
makeEnvironment({ id: "local-1", driver: "local", metadata: { defaultForInstance: false } }),
|
||||
makeEnvironment({ id: "local-2", driver: "local", metadata: null }),
|
||||
];
|
||||
expect(resolveLocalDefaultEnvironmentId(environments)).toBeNull();
|
||||
});
|
||||
|
||||
it("returns null for an empty or missing list", () => {
|
||||
expect(resolveLocalDefaultEnvironmentId([])).toBeNull();
|
||||
expect(resolveLocalDefaultEnvironmentId(null)).toBeNull();
|
||||
expect(resolveLocalDefaultEnvironmentId(undefined)).toBeNull();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,18 +1,128 @@
|
|||
import type { Environment } from "@paperclipai/shared";
|
||||
|
||||
/**
|
||||
* The managed-sandbox-only policy hides the local environment and runs every
|
||||
* agent in the managed sandbox, but no managed sandbox environment is
|
||||
* available. The Test resolution throws this error so the Test surfaces a
|
||||
* fail-closed message, the same way a real run fails closed. The Test must not
|
||||
* fall back to a local host probe, because a host probe would report a result
|
||||
* for a target the real run never uses.
|
||||
*/
|
||||
export class ManagedSandboxUnavailableForTestError extends Error {
|
||||
constructor() {
|
||||
super(
|
||||
"This instance runs agents only in its managed sandbox environment, but no " +
|
||||
"managed sandbox environment is available to test in. Check that the managed " +
|
||||
"sandbox provider is active, then retry the test.",
|
||||
);
|
||||
this.name = "ManagedSandboxUnavailableForTestError";
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Which environment should an adapter "Test" probe?
|
||||
*
|
||||
* Mirrors the server's run-time resolution
|
||||
* (`resolveExecutionWorkspaceEnvironmentId`): the agent's own environment
|
||||
* wins, otherwise the instance default, otherwise none (the server probes
|
||||
* the Paperclip host). Without the instance-default fallback, the Test
|
||||
* button probes the host for agents that rely on the instance default and
|
||||
* fails on commands that only exist inside the default environment — for
|
||||
* example a sandbox image with an extra CLI installed — even though a real
|
||||
* run would have resolved to that environment and succeeded.
|
||||
* The resolution mirrors the server run-time resolution
|
||||
* (`resolveExecutionWorkspaceEnvironmentId`) across all three tiers: the
|
||||
* agent's own environment wins, otherwise the instance default, otherwise the
|
||||
* instance local-default environment. The server always resolves a run to one
|
||||
* of these three tiers, so the Test must probe the same target. Without the
|
||||
* local-default tier the Test would send no environment id and probe the
|
||||
* Paperclip host, even though a real run resolves to the local-default
|
||||
* environment. The two paths must match, so a Test result reflects a real run.
|
||||
*
|
||||
* The managed-sandbox-only policy (`enableManagedSandboxOnly`) redirects a
|
||||
* resolution that lands on the local environment to the managed sandbox
|
||||
* environment, and it fails closed when no managed sandbox exists. This
|
||||
* function mirrors that redirect. A resolution that lands on the local
|
||||
* environment — the local tier, or no tier at all because the policy hides the
|
||||
* local tier — resolves to the managed sandbox environment instead. With no
|
||||
* managed sandbox environment the function throws, so the Test does not probe
|
||||
* the local host that the real run rejects. A resolution that lands on a
|
||||
* non-local environment (an agent or instance default that names an ssh or a
|
||||
* user sandbox) is untouched, because the policy hides local, it does not
|
||||
* forbid other environments.
|
||||
*
|
||||
* The policy hides the local environment from the client, so the resolved id
|
||||
* can name a local row the client cannot see (an agent default that still
|
||||
* points at the hidden local row). The `visibleEnvironmentIds` list closes that
|
||||
* gap: under the policy a resolved id that names no visible environment is the
|
||||
* hidden local row, or a stale reference, so the function redirects it to the
|
||||
* managed sandbox too. Omit the list to skip this visibility check.
|
||||
*/
|
||||
export function resolveAdapterTestEnvironmentId(input: {
|
||||
agentDefaultEnvironmentId: string | null | undefined;
|
||||
instanceDefaultEnvironmentId: string | null | undefined;
|
||||
localDefaultEnvironmentId: string | null | undefined;
|
||||
managedSandboxOnly?: boolean;
|
||||
managedSandboxEnvironmentId?: string | null | undefined;
|
||||
visibleEnvironmentIds?: readonly string[] | null | undefined;
|
||||
}): string | null {
|
||||
return input.agentDefaultEnvironmentId || input.instanceDefaultEnvironmentId || null;
|
||||
const resolved =
|
||||
input.agentDefaultEnvironmentId ||
|
||||
input.instanceDefaultEnvironmentId ||
|
||||
input.localDefaultEnvironmentId ||
|
||||
null;
|
||||
if (input.managedSandboxOnly !== true) {
|
||||
return resolved;
|
||||
}
|
||||
const localDefaultId = input.localDefaultEnvironmentId || null;
|
||||
const resolvedNamesVisibleEnvironment =
|
||||
input.visibleEnvironmentIds == null ||
|
||||
resolved === null ||
|
||||
input.visibleEnvironmentIds.includes(resolved);
|
||||
const landsOnLocal =
|
||||
resolved === null || resolved === localDefaultId || !resolvedNamesVisibleEnvironment;
|
||||
if (!landsOnLocal) {
|
||||
return resolved;
|
||||
}
|
||||
if (!input.managedSandboxEnvironmentId) {
|
||||
throw new ManagedSandboxUnavailableForTestError();
|
||||
}
|
||||
return input.managedSandboxEnvironmentId;
|
||||
}
|
||||
|
||||
/**
|
||||
* Find the active managed sandbox environment id in an environment list. The
|
||||
* platform provisioner stamps the managed environment with
|
||||
* `metadata.managedByPaperclip` (see `isPlatformManagedEnvironment`). The
|
||||
* local-default environment also carries that stamp, so this function excludes
|
||||
* the `local` driver and returns the non-local managed environment. The function
|
||||
* selects only an `active` environment, never an `archived` one, so the Test
|
||||
* resolution matches the server run-time resolver that requires an active
|
||||
* managed sandbox (see `findManagedSandboxEnvironment`). The Test resolution
|
||||
* uses this id as the redirect target under the managed-sandbox-only policy. The
|
||||
* function returns `null` when the list holds no active managed sandbox
|
||||
* environment.
|
||||
*/
|
||||
export function resolveManagedSandboxEnvironmentId(
|
||||
environments: readonly Environment[] | null | undefined,
|
||||
): string | null {
|
||||
if (!environments) return null;
|
||||
const managed = environments.find(
|
||||
(environment) =>
|
||||
environment.driver !== "local" &&
|
||||
environment.status === "active" &&
|
||||
environment.metadata?.managedByPaperclip === true,
|
||||
);
|
||||
return managed?.id ?? null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Find the instance local-default environment id in an environment list. The
|
||||
* server auto-creates one `local` driver environment and stamps it with
|
||||
* `metadata.defaultForInstance: true` (see `ensureLocalEnvironment`). The Test
|
||||
* resolution uses this id as the final tier, so it probes the same environment
|
||||
* a real run resolves to when neither an agent default nor an instance default
|
||||
* is set. The function returns `null` when the list holds no such environment.
|
||||
*/
|
||||
export function resolveLocalDefaultEnvironmentId(
|
||||
environments: readonly Environment[] | null | undefined,
|
||||
): string | null {
|
||||
if (!environments) return null;
|
||||
const local = environments.find(
|
||||
(environment) =>
|
||||
environment.driver === "local" && environment.metadata?.defaultForInstance === true,
|
||||
);
|
||||
return local?.id ?? null;
|
||||
}
|
||||
|
|
|
|||
Loading…
Reference in New Issue