paperclip/evals
Dotta 3db2e6bdd2
feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 8/8 and focuses on end-to-end coverage,
operator docs, evals, and release notes
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack

## Linked Issues or Issue Description

- Related parity reference: #9534
- Problem: The complete stack needs discoverable browser scenarios,
operator guidance, threat modeling, eval coverage, and a parity proof
before merge.
- Proposed solution: Adds MCP user-story and Smoke Lab e2e suites,
docs/evals/release notes, the skill update, and the root e2e driver
script registration.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/07-ui-apps-activation`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for flag audit and e2e/browser acceptance;
Greptile on every PR.

## What Changed

- Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release
notes, the skill update, and the root e2e driver script registration.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.

## Verification

- `pnpm typecheck`
- `node --check scripts/e2e-mcp-user-stories.mjs`
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
--list` — 43 tests discovered
- `git diff pap10341-split/08-e2e-docs
6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes)

## Risks

- Browser suites depend on runtime services and environment setup; this
PR validates discovery locally while QA owns full flag-on/flag-off
execution.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Stack Coordination

- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556#9557#9558#9559#9560#9561#9562#9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-14 15:48:57 -05:00
..
promptfoo feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563) 2026-07-14 15:48:57 -05:00
README.md feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563) 2026-07-14 15:48:57 -05:00

README.md

Paperclip Evals

Eval framework for testing Paperclip agent behaviors across models and prompt versions.

See the evals framework plan for full design rationale.

Quick Start

Prerequisites

pnpm add -g promptfoo

You need an API key for at least one provider. Set one of:

export OPENROUTER_API_KEY=sk-or-...    # OpenRouter (recommended - test multiple models)
export ANTHROPIC_API_KEY=sk-ant-...     # Anthropic direct
export OPENAI_API_KEY=sk-...            # OpenAI direct

Run evals

# Smoke test (default models)
pnpm evals:smoke

# Validate config without provider credentials
cd evals/promptfoo && npx promptfoo@latest validate -c promptfooconfig.yaml

# Or run promptfoo directly
cd evals/promptfoo
promptfoo eval

# Focus only on MCP gateway behavior cases
npx promptfoo@0.103.3 eval -c promptfooconfig.yaml \
  --providers echo \
  --filter-pattern '^mcp_gateway\.' \
  --no-cache \
  --no-progress-bar \
  --no-write

# View results in browser
promptfoo view

What's tested

Phase 0 covers narrow behavior evals for the Paperclip heartbeat skill:

Case Category What it checks
Assignment pickup core Agent picks up todo/in_progress tasks correctly
Progress update core Agent writes useful status comments
Blocked reporting core Agent recognizes and reports blocked state
Approval required governance Agent requests approval instead of acting
Company boundary governance Agent refuses cross-company actions
MCP allowed read tool mcp_gateway Agent records successful gateway calls without unnecessary approval
MCP denied tool mcp_gateway Agent fails closed without retrying or bypassing denied unsafe tools
MCP pending approval mcp_gateway Agent waits on the gateway-created approval path
MCP denied approval mcp_gateway Agent honors rejected or unapproved tool actions
MCP rate limit mcp_gateway Agent backs off without crashing or busy-looping
MCP missing credential mcp_gateway Agent blocks on credential repair without leaking or inventing secrets
MCP revoked session mcp_gateway Agent stops using stale gateway tokens and avoids raw upstream fallback
MCP header forwarding mcp_gateway Agent reports forwarded transport/credential headers from redacted audit evidence
MCP named target mcp_gateway Agent uses the exact on-demand named gateway tool rather than an ambiguous upstream name
MCP elicitation mcp_gateway Agent asks the human/board for missing input instead of fabricating it
MCP approved target drift mcp_gateway Agent treats changed catalog/schema/credential snapshots as stale approval
No work exit core Agent exits cleanly with no assignments
Checkout before work core Agent always checks out before modifying
409 conflict handling core Agent stops on 409, picks different task
Memory provider binding phase5_memory Agent honors agent override before company default
Memory provenance audit phase5_memory Agent preserves inspectable source and operation records
Memory hook cost/trust phase5_memory Agent keeps memory hook cost attribution and source trust visible
Board command work objects phase5_control_surface Chat-like board commands create auditable work objects

Phase 5 memory/control-surface prompt evals should be paired with deterministic server/shared tests for:

  • memory provider resolution order: agent override, then company default
  • memory operation audit rows including company, agent, issue, run, provider, source, and cost references
  • hook-delivered memory payloads preserving source trust and cost attribution fields
  • board command/chat-like routes creating auditable issues, comments, documents, approvals, or work products

Adding new cases

  1. Add a YAML file to evals/promptfoo/tests/
  2. Follow the existing case format (see core.yaml for reference)
  3. Run promptfoo eval to test

Phases

  • Phase 0 (current): Promptfoo bootstrap - narrow behavior evals with deterministic assertions
  • Phase 1: TypeScript eval harness with seeded scenarios and hard checks
  • Phase 2: Pairwise and rubric scoring layer
  • Phase 3: Efficiency metrics integration
  • Phase 4: Production-case ingestion