|
…
|
||
|---|---|---|
| .. | ||
| promptfoo | ||
| README.md | ||
README.md
Paperclip Evals
Eval framework for testing Paperclip agent behaviors across models and prompt versions.
See the evals framework plan for full design rationale.
Quick Start
Prerequisites
pnpm add -g promptfoo
You need an API key for at least one provider. Set one of:
export OPENROUTER_API_KEY=sk-or-... # OpenRouter (recommended - test multiple models)
export ANTHROPIC_API_KEY=sk-ant-... # Anthropic direct
export OPENAI_API_KEY=sk-... # OpenAI direct
Run evals
# Smoke test (default models)
pnpm evals:smoke
# Validate config without provider credentials
cd evals/promptfoo && npx promptfoo@latest validate -c promptfooconfig.yaml
# Or run promptfoo directly
cd evals/promptfoo
promptfoo eval
# Focus only on MCP gateway behavior cases
npx promptfoo@0.103.3 eval -c promptfooconfig.yaml \
--providers echo \
--filter-pattern '^mcp_gateway\.' \
--no-cache \
--no-progress-bar \
--no-write
# View results in browser
promptfoo view
What's tested
Phase 0 covers narrow behavior evals for the Paperclip heartbeat skill:
| Case | Category | What it checks |
|---|---|---|
| Assignment pickup | core |
Agent picks up todo/in_progress tasks correctly |
| Progress update | core |
Agent writes useful status comments |
| Blocked reporting | core |
Agent recognizes and reports blocked state |
| Approval required | governance |
Agent requests approval instead of acting |
| Company boundary | governance |
Agent refuses cross-company actions |
| MCP allowed read tool | mcp_gateway |
Agent records successful gateway calls without unnecessary approval |
| MCP denied tool | mcp_gateway |
Agent fails closed without retrying or bypassing denied unsafe tools |
| MCP pending approval | mcp_gateway |
Agent waits on the gateway-created approval path |
| MCP denied approval | mcp_gateway |
Agent honors rejected or unapproved tool actions |
| MCP rate limit | mcp_gateway |
Agent backs off without crashing or busy-looping |
| MCP missing credential | mcp_gateway |
Agent blocks on credential repair without leaking or inventing secrets |
| MCP revoked session | mcp_gateway |
Agent stops using stale gateway tokens and avoids raw upstream fallback |
| MCP header forwarding | mcp_gateway |
Agent reports forwarded transport/credential headers from redacted audit evidence |
| MCP named target | mcp_gateway |
Agent uses the exact on-demand named gateway tool rather than an ambiguous upstream name |
| MCP elicitation | mcp_gateway |
Agent asks the human/board for missing input instead of fabricating it |
| MCP approved target drift | mcp_gateway |
Agent treats changed catalog/schema/credential snapshots as stale approval |
| No work exit | core |
Agent exits cleanly with no assignments |
| Checkout before work | core |
Agent always checks out before modifying |
| 409 conflict handling | core |
Agent stops on 409, picks different task |
| Memory provider binding | phase5_memory |
Agent honors agent override before company default |
| Memory provenance audit | phase5_memory |
Agent preserves inspectable source and operation records |
| Memory hook cost/trust | phase5_memory |
Agent keeps memory hook cost attribution and source trust visible |
| Board command work objects | phase5_control_surface |
Chat-like board commands create auditable work objects |
Phase 5 memory/control-surface prompt evals should be paired with deterministic server/shared tests for:
- memory provider resolution order: agent override, then company default
- memory operation audit rows including company, agent, issue, run, provider, source, and cost references
- hook-delivered memory payloads preserving source trust and cost attribution fields
- board command/chat-like routes creating auditable issues, comments, documents, approvals, or work products
Adding new cases
- Add a YAML file to
evals/promptfoo/tests/ - Follow the existing case format (see
core.yamlfor reference) - Run
promptfoo evalto test
Phases
- Phase 0 (current): Promptfoo bootstrap - narrow behavior evals with deterministic assertions
- Phase 1: TypeScript eval harness with seeded scenarios and hard checks
- Phase 2: Pairwise and rubric scoring layer
- Phase 3: Efficiency metrics integration
- Phase 4: Production-case ingestion