* feat: add better params to working representation fetch in SDKs, return messages when added * fix: working representation routes now accepting all parameters properly, with tests * feat: add metadata/config fields to SDK objects where viable * fix: tests * feat: refactor SDKs to use representation config; [TEMP STAINLESS BUILD] update API * feat: add representation object to sdks * fix: use stainless sdk on branch * fix: update TypeScript SDK tsconfig to use node16 module resolution * fix: add isolatedModules = true to tsconfig * fix: lol * chore: coderabbit review * feat: make delete session real * feat: add observations routes with delete endpoints for documents. make session deletion real. * chore: type cleanup * fix: tests * chore: coderabbit review * fix: namespace by workspace * feat: add ability to customize messages_per_summary at both workspace and session level * chore: tests for summary config * chore: coderabbit cleanup * feat: make session and workspace config totally customizeable * feat: add search by peer knowledge (#250) * feat: search by peer perspective * fix: enforce workspace in filters, make messages distinct in join * fix: batch and merge migration steps * fix: add refresh, add config to workspace, add refresh function, make fields readonly * fix: search distinct * fix: merge migrations * fix: merge migrations * fix: batch deletions, improve comments, limit consolidate dream to 100 docs at a time, auth on observations routes * chore: review * chore: coderabbit * chore: review * chore: broken comment * feat: add set peer card route to API * feat: create advanced configuration parameters with message>session>workspace hierarchy * [wip] build unified testing harness * chore: lint * fix: cache invalidation, naming things, etc * feat: longmem tests * chore: peer config refactor * feat: consolidate dream working, refactor representation * fix: Various CR Comment Fixes * feat: Allow configurable Redis port for harness instances and update cleanup methods to be asynchronous. * feat: agentic ingestion task!!! * feat: agentic deriver * feat: dialectic agent and dreamer agent * chore: browbeat tests into passing * fix: nits * chore: remove old code, update config files * fix: simplify deriver * feat: dialectic agent prompt updates, re-introduce non_agent deriver, eval tweaks * feat: fast deriver, dreamer, then dialectic * fix: tweaks across the board * feat: add baseline tests * feat: truncation in tools and client, tweaks for evals * feat: add locomo, fix longmem judge!!! * fix: locomo f1 is trash, use llm judge * feat: trace creation * feat: add first draft of obex benchmark, fix embedding model, fix locomo methodology * fix: locomo session-optimized, better logging of cache usage and better cache usage * chore: use openrouter for baselines * fix: add test for merge migration * chore: opus-powered cleanup * fix: add config for vllm, better client * chore: clean up clients.py a bit * chore: move magic numbers to config, add tests for agent tools * fix: wrong mock in dialectic tests, make ToolContext a dataclass * feat: tweak prompts, make deriver explicit-only * feat: more prompt & tool tweaks * chore: more tweaks * feat: dream with subagents * fix: make dream trigger override scheduled, play around with dream agents * chore: cleanup deriver * chore: cleanup dialectic * chore: cleanup orchestrator * chore: comment out dream stuff, WIPing * fix: inc temp on retry, typechecking * feat: tweak dreaming * feat: contradiction obs * Add dream trees * chore: preserve reasoning_details from openrouter in client * fix: get_observation_context correct params * fix: use correct message id in tool * chore: cleanup longmem runner * chore: clean up tests, remove dream tests for now as rearchitecting around trees * chore: update stainless deps * Update threholding mechanism * chore: pre-commit hooks whitespace * chore: clean up types * feat: add explicit bench * fix: address additional basepyright issues * fix: adding logging as a fixture on honcho_llm_call and supporting dialectic loging. (#305) * fix: lock on db for tool calls * chore: clean up experimental derivers * chore: coderabbit review cleanup * feat: add streaming support to agentic dialectic * feat: prometheus token tracking for deriver and dialectic * fix: self-loops for isolated nodes * chore: PascalCase for prometheus parameter typing * feat: add reasoning levels to dialectic agent * chore: delete old file, add new fake env vars in unittest.yml * fix: all fields needed for dialectic reasoning level configs * feat: track dreaming usage in prometheus * chore: Create backwards compatabile conclusion and queue endpoints * fix: remove redundant try-catch, add trace label, move .limit to end of statement * fix: remove vignettes (for now), review fixes, remove merge migration, config cleanup * chore: code review / cleanup * chore: merge fixes * chore: clean up, remove reasoning_focus, reintroduce peer cards in dreamers * chore: code rabbit nitpicks * fix: add unique index for pending dreams in queue * fix: revert removal of surprisal in dreamer config --------- Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com> Co-authored-by: 3un01a <3un01a@plasticlabs.ai> Co-authored-by: 3un01a <3un01a.labs@gmail.com> Co-authored-by: ajspig <46900795+ajspig@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| README.md | ||
| __init__.py | ||
| conftest.py | ||
| test_deriver_processing.py | ||
| test_queue_operations.py | ||
| test_queue_processing.py | ||
| test_representation_crud.py | ||
README.md
Deriver Testing
This directory contains tests for the deriver system, which handles background processing of messages to extract insights and update working representations.
Structure
conftest.py- Shared fixtures for deriver testingtest_queue_operations.py- Tests for basic queue operationstest_deriver_processing.py- Tests for deriver processing logictest_queue_processing.py- Tests for queue manager and work unit processing
Key Fixtures
Database Fixtures
sample_session_with_peers- Creates a session with multiple peers having different observation configurationssample_messages- Creates sample messages for testingsample_queue_items- Creates queue items with various payload types (representation, summary)
Queue Fixtures
create_queue_payload- Helper to create queue payloads for testingadd_queue_items- Helper to add queue items to the databasecreate_active_queue_session- Helper to create active queue sessions for work unit tracking
Mocking Fixtures
mock_critical_analysis_call- Mocks the critical analysis LLM callmock_queue_manager- Mocks the queue manager for testingmock_representation_manager- Mocks the representation manager operations
Testing Patterns
Creating Queue Items
# Create representation payloads
payload = create_queue_payload(
message=message,
task_type="representation",
observer=observer_peer.name,
observed=message.peer_name
)
# Add to queue
queue_items = await add_queue_items([payload], session.id)
Testing Work Units
# Create a work unit
work_unit = WorkUnit(
session_id=session.id,
task_type="representation",
observer=observer,
observed=observed
)
# Test string representation
assert str(work_unit) == f"({session.id}, {observed.name}, {observer.name}, representation)"