* chore: 3.0 honcho and 2.0 sdks changelog fix: use PeerContextResponse in peer.ts * chore: move docs to /v3/, build SDKs * chore: code review * feat: [WIP] migrate away from stainless in typescript sdk * chore: move api from /v2/ to /v3/ * feat: no-stainless typescript with real tests * feat: migrate python sdk off of stainless * feat: clean typescript sdk * chore: add tests for ts http client * fix: rewrite entire python sdk in new format, update typescript sdk to use `configuration` not `config` for consistency with API * fix: clean up SDKs, synchronize * chore: update sdk examples * chore: update OpenAPI documentation and SDK examples to reflect changes * fix: better test * fix: install deps in test runner, improve robustness of streaming in sdk, coderabbit nits * fix: standardize around camelCase in TS SDK * refactor: update configuration handling in SDKs to use typed models for workspace, session, and peer configurations * docs: clarify queue status usage and remove polling methods from SDKs add claude skills for migrations * chore: fix links in docs * feat: add deriver flush mode to bypass batch token threshold - Introduced `is_deriver_flush_enabled` function to check if flush mode is active. - Updated `QueueManager` to conditionally apply batch token thresholds based on flush mode. - Enhanced `UnifiedTestExecutor` to enable flush mode via Redis. - Added `flush` parameter to test cases to facilitate testing of flush mode behavior. - Updated various test cases to utilize the new flush functionality. * feat: implement schedule_dream functionality in SDKs, use in unified test runner - Added `schedule_dream` method to both Python and TypeScript SDKs for scheduling dream tasks. - Updated HTTP routes to include endpoint for scheduling dreams. - Enhanced test runner to utilize the new `schedule_dream` method for scheduling actions. - Updated TypeScript client to support the new scheduling functionality with appropriate parameters. * feat: update single deriver task to support multiple observers - Changed the `observer` parameter to `observers` as a list in multiple functions across the deriver module. - Updated the processing logic to handle multiple observers for representation tasks. - Adjusted related payload and queue management functions to accommodate the new observers structure. - Modified tests to reflect changes in the representation task handling and ensure proper functionality. * refactor: update enqueue tests to support deduplication of queue items with multiple observers - Modified tests in `test_enqueue.py` to reflect changes in the queue item structure, where each message now results in a single queue item containing a list of observers. - Updated assertions to validate that the `observers` field correctly includes all relevant peers, ensuring proper functionality of the deduplication logic. - Removed redundant payload matching logic to streamline test cases and improve clarity. * fix: add backwards compatibility for representation work unit keys and payload observers * feat: update dialectic configuration and introduce cost calculator - Adjusted LLM and dialectic settings in `.env.template`, `config.toml.example`, and `src/config.py` to reduce maximum tool output characters and session history tokens for cost efficiency. - Implemented a new `dialectic_cost_calculator.py` script to estimate costs based on reasoning levels and model pricing. - Enhanced `DialecticAgent` to utilize minimal tools and adjusted output token settings based on reasoning level to optimize performance and reduce costs. * feat: add reasoning level to chat input in unified test runner - Enhanced the `UnifiedTestExecutor` to include a `reasoning_level` parameter in the chat method call. - Updated the `QueryAction` model to support the new `reasoning_level` attribute, allowing for more nuanced chat interactions. * feat: run deriver once for multiple observers (#335) * feat: update single deriver task to support multiple observers - Changed the `observer` parameter to `observers` as a list in multiple functions across the deriver module. - Updated the processing logic to handle multiple observers for representation tasks. - Adjusted related payload and queue management functions to accommodate the new observers structure. - Modified tests to reflect changes in the representation task handling and ensure proper functionality. * refactor: update enqueue tests to support deduplication of queue items with multiple observers - Modified tests in `test_enqueue.py` to reflect changes in the queue item structure, where each message now results in a single queue item containing a list of observers. - Updated assertions to validate that the `observers` field correctly includes all relevant peers, ensuring proper functionality of the deduplication logic. - Removed redundant payload matching logic to streamline test cases and improve clarity. * fix: add backwards compatibility for representation work unit keys and payload observers * feat: refactor benchmark runners to share common functionality - Introduced a new `runner_common.py` module containing shared utilities for benchmark test runners, including common argument parsing, client creation, and queue management. - Updated `BEAMRunner`, `LoCoMoRunner`, and `LongMemEvalRunner` to inherit from `RunnerMixin`, leveraging shared functionality for metrics collection and logging. - Added `reasoning_level` and `redis_url` parameters to runner constructors for enhanced configuration. - Streamlined argument parsing by utilizing `add_common_arguments` for shared command-line options across all runners. * fix: update last_user_message handling to use message content instead of ID * fix: standardize config vs configuration --------- Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| test_cases | ||
| README.md | ||
| run.py | ||
| runner.py | ||
| schema.py | ||
README.md
Unified Honcho Test System
This system allows for defining comprehensive, step-based tests for Honcho in a unified JSON format. It supports testing configuration hierarchy, multi-turn interactions, and complex assertions including LLM-as-a-judge.
Running Tests
# Run all tests in the test_cases directory
python -m tests.unified.run
# Run a specific test file
python -m tests.unified.run --test-dir tests/unified/test_cases
Test Schema
Tests are defined in JSON files. A test definition consists of a name, optional description, and a list of steps.
Structure
{
"name": "my_test",
"workspace_config": { ... },
"steps": [
{ "step_type": "..." },
...
]
}
Actions
-
Configuration:
set_workspace_config: Update workspace settings.set_session_config: Update session settings.
-
Interaction:
create_session: Create a new session, optionally with peers and config.add_message: Add a single message.add_messages: Add multiple messages.
-
Waiting:
wait: Wait for duration or "queue_empty".
-
Querying & Assertions:
query: Perform an action and assert on the result.target: "chat", "get_context", "get_peer_card", "get_representation"
Assertions
llm_judge: Use Claude to evaluate the result against a natural language prompt.contains/not_contains: Substring matching.exact_match: Strict equality.json_match: specific key-value checks.
Example
{
"name": "demo_config_flow",
"steps": [
{
"step_type": "create_session",
"session_id": "s1",
"peer_configs": {
"user": { "observe_me": true },
"agent": { "observe_others": true }
}
},
{
"step_type": "add_message",
"session_id": "s1",
"peer_id": "user",
"content": "My name is Alice."
},
{
"step_type": "wait",
"target": "queue_empty"
},
{
"step_type": "query",
"target": "chat",
"peer_id": "agent",
"session_id": "s1",
"input": "Who am I?",
"assertions": [
{
"assertion_type": "contains",
"text": "Alice"
}
]
}
]
}