# Unified Honcho Test System This system allows for defining comprehensive, step-based tests for Honcho in a unified JSON format. It supports testing configuration hierarchy, multi-turn interactions, and complex assertions including LLM-as-a-judge. ## Running Tests ```bash # Run all tests in the test_cases directory python -m tests.unified.run # Run a specific test file python -m tests.unified.run --test-dir tests/unified/test_cases ``` ## Test Schema Tests are defined in JSON files. A test definition consists of a name, optional description, and a list of steps. ### Structure ```json { "name": "my_test", "workspace_config": { ... }, "steps": [ { "step_type": "..." }, ... ] } ``` ### Actions 1. **Configuration**: * `set_workspace_config`: Update workspace settings. * `set_session_config`: Update session settings. 2. **Interaction**: * `create_session`: Create a new session, optionally with peers and config. * `add_message`: Add a single message. * `add_messages`: Add multiple messages. * `create_scope`: Create a scope and optionally add member sessions. Add the sessions *before* the messages you want in scope — membership only affects messages ingested after a session joins. 3. **Waiting**: * `wait`: Wait for duration or "queue_empty". 4. **Querying & Assertions**: * `query`: Perform an action and assert on the result. * `target`: "chat", "get_context", "get_peer_card", "get_representation", "workspace_chat" * `scope`: confine the read to a scope (or, for chat/representation/ workspace_chat, to the union of several). Valid for "chat", "get_representation", "get_context", and "workspace_chat"; get_context takes a single scope and requires `observed_peer_id`. ### Raw HTTP vs the SDK Most steps drive the Honcho Python SDK. `create_scope` and scoped `chat` / `get_representation` / `get_context` queries go over raw HTTP instead, because the published SDK trails the API and exposes neither. Scoped `workspace_chat` uses the SDK `scope` argument. Calling the API directly also tests the contract the SDK is generated from, so a wrong status code or response shape surfaces here rather than being masked by client-side validation. ### Assertions * `llm_judge`: Use Claude to evaluate the result against a natural language prompt. * `contains` / `not_contains`: Substring matching. * `exact_match`: Strict equality. * `json_match`: specific key-value checks. ## Example ```json { "name": "demo_config_flow", "steps": [ { "step_type": "create_session", "session_id": "s1", "peer_configs": { "user": { "observe_me": true }, "agent": { "observe_others": true } } }, { "step_type": "add_message", "session_id": "s1", "peer_id": "user", "content": "My name is Alice." }, { "step_type": "wait", "target": "queue_empty" }, { "step_type": "query", "target": "chat", "peer_id": "agent", "session_id": "s1", "input": "Who am I?", "assertions": [ { "assertion_type": "contains", "text": "Alice" } ] } ] } ```