44 KiB
44 KiB
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog and this project adheres to Semantic Versioning.
[3.0.11] - 2026-06-24
Added
api_request_duration_secondsPrometheus histogram tracking per-route request latency, labeled by method and endpoint (#837)- LLM
provider_paramspassthroughs (extra_body/extra_headers/extra_query) are now forwarded to the underlying provider transport across all backends, with shape validation that rejects non-mapping values (#821) structured_output_modemodel-config option to usejson_objectmode for OpenAI-compatible providers that lack native Structured Outputs support (used by the deriver) (#820)- OpenRouter app-attribution headers (
HTTP-Referer/X-Openrouter-Title) are now sent on OpenAI-compatible clients when the configured base URL is OpenRouter, so requests are attributed to "Honcho" in OpenRouter's dashboard (#805) - Langfuse traces are now tagged with user and session IDs for easier trace filtering (#814)
DERIVER_REPRESENTATION_BATCH_MAX_AGE_SECONDS(default 1800s) lets sub-threshold representation work units flush once their oldest unprocessed queue item ages out. Set it to0to keep the legacy behavior where sub-threshold tails wait indefinitely unlessDERIVER_FLUSH_ENABLED=true(#826)- Conclusion responses now include a
levelfield (explicit,deductive,inductive,contradiction); list/query endpoints support filtering bylevelviafilters, with reserved filter keys protected from being overridden by user-supplied filters (#851)
Changed
- Peer-scoped JWTs now get read-only access to the sessions their peer is an active member of (session context, summaries, peers, their own per-session config, search, and message reads). Session-scoped JWTs remain confined to their session and cannot reach peer routes (#679)
- Compacted Honcho's log output, with guarded ms/s metric formatting that falls back to a plain string for non-numeric values (#836)
- Sentry now drops noisy infra/scrape transactions: the reconciler opens a transaction only once a batch has rows (idle cycles emit none), and a
traces_samplerreturns0.0for/metrics,/health,/openapi.json,/docs,/redoc, and the deriver metrics server.SENTRY.TRACES_SAMPLE_RATEstill governs real traffic (#834)
Fixed
- Peer- and session-scoped JWTs were effectively workspace-scoped: authorization walked the route's declared scope and fell through to a workspace match, so a
{w, p: alice}token could act on any peer in the workspace. JWTs are now authorized by their narrowest claim and never widen to workspace access (#679) - The keys API now rejects creating a peer- or session-scoped key without a workspace. Such keys were minted successfully but failed verification on every request (#679)
- Agent-supplied observation IDs carrying the display-format
id:prefix are now normalized (prefix and trailing whitespace stripped) beforesource_idsare stored and onget_reasoning_chainlookups, fixing corrupted provenance links and broken reasoning-chain traversal (#795) - Fixed a
create_treekeyword-argument mismatch in the Dreamer's surprisal tree construction (#749) - Providers that omit output-token counts (observed with Gemini on tool-loop completions) returned
output_tokens=None, which raised a Pydantic validation error that aborted the call and crashed the Dreamer's induction phase before inductive conclusions were persisted.Noneis now coerced to0so token accounting degrades gracefully (#809) - Document creation now performs exact (case-insensitive, whitespace-trimmed) content deduplication before the existing semantic dedup step: exact duplicates within a batch collapse to a single insert, and an exact match against a live document reinforces it (atomic
times_derivedincrement) instead of creating a new row (#861)
[3.0.10] - 2026-06-15
Added
- Messages are now embedded via a background task rather than blocking API request
- Read-only DB session mode (
get_read_db/tracked_db(..., read_only=True)) so reads don't hold a transaction open across the work CORS_ORIGINSenv var to configure CORS allowed origins without editing source; defaults match the prior hardcoded list, so self-hosted deployments behind custom domains can whitelist their frontend (#697)scripts/generate_jwt.py— utility for minting scoped or admin Honcho JWTs (--admin,--workspace/--peer/--session,--expireswith human-friendly durations,--print-only) without calling the keys API (#757)STALE_WORK_UNIT_CLEANUP_INTERVAL_SECONDS(default 60s) — minimum jittered spacing between deriver stale-work-unit cleanup runs, so cleanup no longer runs on every seconds-scale poll (0.0keeps the legacy every-poll behavior) (#773)
Changed
- Optimized the deriver and dreamer prompt cache prefixes to improve prompt-cache hit rates (#806)
Fixed
times_derivedis now properly reinforced when a duplicate conclusion is detected. It had been pinned at 1 for nearly every conclusion (the reject-new branch dropped the increment and the new-wins branch reset the count to 1), soORDER BY times_derived DESCfell back to arbitrary heap order and froze stale conclusions to the front of injected context. Reinforcement is now an atomic increment and both most-derived queries gained acreated_at DESCrecency tiebreaker (#768)- Webhook creation now correctly rejects private/internal IP addresses (#793)
[3.0.9] - 2026-06-02
Changed
- Connection acquisition is now a single attempt with no server-side retry, on a vanilla
AsyncSession. A newDB_CONNECT_TIMEOUT_SECONDS(default 2s) bounds the attempt so a saturated or unreachable pooler fails fast instead of holding a client connection open to re-knock. A saturated DB now surfaces to the caller — the API returns an error and the deriver backs off and retries on a later poll — which lets the pooler drain rather than amplifying saturation.
Added
- Deriver poll jitter so instances that start together don't poll in lockstep:
DERIVER_POLLING_STARTUP_JITTER_SECONDS(random delay before the first poll, default 30s) andDERIVER_POLLING_JITTER_RATIO(±fraction applied to every poll sleep, default 0.5). Both disable at0.0; the underlying backoff schedule is unchanged.
Removed
- Reverted the connection-checkout retry and
HonchoAsyncSessioncustom session introduced in 3.0.8. Removed theDB_CONNECTION_RETRY_ENABLED/DB_CONNECTION_RETRY_MAX_DELAY_SECONDS/DB_CONNECTION_RETRY_BACKOFF_INITIAL_SECONDS/DB_CONNECTION_RETRY_BACKOFF_MAX_SECONDSsettings, thedb_connection_acquisitions{outcome=...}Prometheus counter, and thedb.pool.acquireSentry span. Alerting built ondb_connection_acquisitionsshould migrate todb_pool_connections/db_queries_in_flight.
[3.0.8] - 2026-06-01
Added
- Connection-checkout retry with bounded exponential backoff (tenacity) on
get_db/tracked_db: transient transaction-pooler (Supavisor) rejections — SQLAlchemyTimeoutErrorandOperationalError— now retry with backoff instead of surfacing as 500s under client-connection saturation. Gated byDB_CONNECTION_RETRY_ENABLEDwith configurable delay/backoff knobs; ~10s default budget (#758) HonchoAsyncSession— a lazyAsyncSessionthat checks out its pooled connection (with retry) on the first DB-touching call rather than at construction. Request handlers doing non-DB work (embedding, file, LLM) before their first query no longer pin a pooler connection across it. Only the checkout is retried; the statement still runs exactly once, so writes are never duplicated (#758)- Adaptive deriver queue polling: the poll interval backs off when the queue is idle or erroring (base → max, doubling each cycle) and snaps back to base the moment work is claimed, cutting steady-state query load against the DB. Gated by
DERIVER_POLLING_BACKOFF_ENABLEDwith configurable max/multiplier (#758) - New Prometheus
db_pool_connectionsgauge (checked_out / checked_in / size / overflow), labeledapi|deriver, registered in both the API lifespan and the deriver metrics server (#758) - New Prometheus
db_connection_acquisitions{outcome=ok|retried|exhausted}counter — the alertable early-warning signal that connection checkouts are retrying through pooler rejection, before requests start failing (#758) - New Prometheus
db_queries_in_flightgauge — statements actually executing on the wire (via SQLAlchemy cursor-execute events). Paired withchecked_out, the gap reveals connections held but parked (the "idle in transaction during an external call" antipattern). Gated onMETRICS.ENABLEDfor zero overhead when off (#758) - Explicit
SqlalchemyIntegrationin both the API and deriver Sentry inits; connection acquisition wrapped in adb.pool.acquirespan with live pool stats captured on retry exhaustion (#758)
Changed
- Default
POOL_TIMEOUTlowered to 5s, with validation that it stays under the connection-retry budget when a pooled (non-null)POOL_CLASSis configured;config.toml.exampleand the v2/v3 configuration docs updated to match (#758) HonchoAsyncSessionwraps every DB-touching session method (execute / scalar / scalars / flush / merge / refresh / commit / get / get_one / stream / stream_scalars / delete) so the lazy-checkout-with-retry guarantee has no holes; the acquired flag resets onclose()/reset()so a reused session re-acquires on next use (#758)
Fixed
- Roll the session back on a retryable checkout failure before retrying — a failed autobegin could otherwise leave it pending-rollback, making the next connection attempt raise instead of cleanly re-checking-out (#758)
- Guard
DBPoolCollector.collect()so a pool-read/import hiccup can't raise and abort the entire/metricsscrape (Prometheus drops all metrics if any collector raises) (#758) - Clamp the pool overflow gauge to ≥ 0 (it could report negative before the pool fills) (#758)
- Removed a double-sleep in the deriver idle poll so the backoff cap is a true cap rather than 2× (#758)
[3.0.7] - 2026-05-21
Added
- New
src/llm/module as the single owner of provider runtime: clients, backends, history adapters, tool loop, request builder, credentials, and caching policy (#459) AttemptPlandataclass captures per-retry provider selection (client, model, reasoning_effort, thinking_budget_tokens, selected_config) and pins it across stream-final retries so streaming doesn't bounce back to primary after the tool loop has settled on fallback (#459)- Gemini JSON-schema sanitizer for
function_declarations— strips keywords Gemini's validator rejects (additionalProperties,allOf, etc.) while preserving semantics for all other backends (#459) - Dreamer specialists derive
effective_max_tokensfrommodel_config.max_output_tokenswith a per-specialist default fallback (#459) - New cloudevent
LLMCallCompletedEvent(llm.call.completed) fires once per provider hit with full cost-attribution context: transport/provider_label, model, token counts with cache breakdown, finish_reason, outcome,is_final_attempt, retry/fallback state, duration, tool-call shape, streaming flag, and agent correlation (run_id+ iteration). Includes aCallPurposeclosed enum (deriver.representation,dialectic.answer,dream.deduction|induction,summary.short|long) (#637) RepresentationCompletedEventnow carriestotal_input_tokensfor full-trace cost attribution (#637)- Per-emitter
honcho_versioninjection on all CloudEvents plus emitter health metrics (#637) TelemetrySettings.HIGH_VOLUME_SAMPLE_RATE(default 1.0) — deterministic per-run_idsampler so an entire agent trace is kept or dropped together; aggregate envelopes bypass the sampler (#637)- Deriver custom instructions: per-workspace/peer guidance threaded into the deriver prompt with a
MAX_CUSTOM_INSTRUCTIONS_TOKENSbudget (default 2000); deriverMAX_INPUT_TOKENSraised 23000 → 25000 to make room (#609) - Configurable embedding dimensions:
EMBEDDING_MODEL_CONFIG__DIMENSIONS_MODE(auto/always/never) controls whether the OpenAIdimensions=parameter is forwarded;auto(default) sends it when the operator explicitly setEMBEDDING_VECTOR_DIMENSIONSand the model is not on the known-rejecting allowlist (#678) - New
honcho-clipackage — Python CLI for inspecting and managing peers, sessions, and configuration against a Honcho deployment (#424) HONCHO_API_URLenv var support in the MCP Worker, enabling self-hosted Honcho deployments to point the Worker at their own instance instead ofhttps://api.honcho.dev(#575)- API ID
max_lengthincreased from 100 to 512 acrossWorkspaceCreate,PeerCreate, andSessionCreateto align the API contract with the underlying DB schema (#684) - Regression tests covering fallback-config thinking-param reach, provider_params → extra_params boundary, OpenAI reasoning-model parameter routing, Gemini blocked finish_reason handling, and fail-fast
max_tool_iterationsvalidation (#459)
Changed
- All LLM orchestration moved out of
src/utils/clients.pyintosrc/llm/with modules split by responsibility (api, executor, tool_loop, runtime, registry, conversation, request_builder, credentials, caching, backends, history_adapters) (#459) - Default
ModelConfigfactories (deriver, summary, dreamer specialists, dialectic levels) normalized toopenai/gpt-5.4-miniwith no extra parameters set by default; operators add transport/thinking overrides explicitly (#459) - OpenAI reasoning-model routing widened via
_uses_max_completion_tokensheuristic coveringgpt-5.xando1/o3/o4— these models receivemax_completion_tokensinstead ofmax_tokens(#459) - Override client factories switched from unbounded
@cacheto@lru_cache(maxsize=128)for predictable memory growth on long-running processes (#459) get_backendnow delegates toclient_for_model_config, so the live-test path and production path share one missing-API-key validation (#459)- Blocked Gemini responses (
SAFETY,RECITATION,PROHIBITED_CONTENT,BLOCKLIST) raiseLLMErrorin the streaming path too (previously only the non-streaming path), ensuring retry/fallback logic fires uniformly (#459) - Transport-change env overrides now strip transport-specific thinking params (thinking_budget_tokens vs. reasoning_effort) during config merge, including at the dialectic-level merge, so switching from Anthropic → OpenAI doesn't leave orphaned Anthropic-only params that the OpenAI backend would reject (#459)
max_tool_iterationsout-of-range inputs now raiseValidationExceptioninstead of being silently clamped (#459)- Public API schemas (
WorkspaceCreate,PeerCreate,SessionCreate) and SDK validation (api_types.py,validation.ts) accept IDs up to 512 chars (was 100) (#684) - Peer card prompts reframed as stable identity markers (replaces the prior "biographical/profile facts" language). Induction specialist is now opted out of peer card writes (
can_update_peer_card = False) so only deduction touches the card (#686) - Vector store queries no longer fetch embedding vectors — only document metadata is returned, reducing payload size and DB load (pgvector, lancedb, turbopuffer) (#682)
- Langfuse trace metadata now includes
namespace,model, andproviderso traces can be filtered by deployment slice (#565) - Deriver: model-aware tokenizer (replaces the previously hardcoded encoding) and explicit guard on empty message content (#647)
- Dialectic level defaults now merge correctly with per-level overrides in
src/config(DEV-1733) (#656) - Default dialectic tool choice switched from forced/required to
auto(#630) - Vector sync given a substantial retry budget to tolerate transient embedding provider outages (#604)
AgentToolConclusionsDeletedEventpayload now carrieslevelsfor parity with the rest of the conclusion event surface (#612)- Turbopuffer vector store:
InternalServerErrorcaught and surfaced as a warning rather than a hard failure; unusedupsert_with_retryandVectorUpsertResultremoved; explicit silent and explicit-error paths for vector DB server errors (#561) - Troubleshooting docs updated to reflect nested-env-var form for per-component thinking-budget overrides (#459)
- README refresh (#681)
- CLAUDE.md refreshed against the current
src/layout (#680)
Fixed
- Fallback
ModelConfigtemperature andthinking_budget_tokensreach the backend on the final retry — previously the primary's values were pre-populated into caller kwargs early and clobbered fallback values viaeffective_config_for_call(update=...)(#459) - Stream-final retries pin to the
AttemptPlanthat succeeded rather than re-running provider selection through the outercurrent_attemptContextVar (which could roll streaming back to primary after the tool loop had already switched to fallback) (#459) - OpenAI structured-output calls continue to use
chat.completions.parse()with strict schema enforcement, while tool-calling paths usechat.completions.create()withoutstrict:Truefor broader proxy compatibility (OpenRouter, vLLM, Ollama) (#459) - Gemini
cached_contentreuse keys now includesystem_instructionandtool_configso cache hits don't cross configurations that differ only in those fields (#459) - Removed strict parameter validation for thinking params on Anthropic and OpenAI transports — was rejecting valid per-transport configs (#686)
reversequery parameter is now honored on the v3 workspace list (POST /v3/workspaces/list), peer list (POST /v3/workspaces/{workspace_id}/peers/list), workspace-scoped session list (POST /v3/workspaces/{workspace_id}/sessions/list), and peer-scoped session list (POST /v3/workspaces/{workspace_id}/peers/{peer_id}/sessions). Honcho SDKs at 2.1.0+ were already sendingreverse=truefor these routes but the server silently ignored it. Ties oncreated_atnow fall back to the internal nanoididso ordering remains stable across pages (#685)- LLM client factories now receive
base_urlfromLLMSettingsfor default providers — previously the override path honoredbase_urlbut the default path didn't, so operators pointing at OpenAI-compatible proxies viaLLM__OPENAI_BASE_URLwere ignored (#643, fixes #641) - Internal N+1 query in dialectic agent tool execution (DEV-1721) — collapsed per-iteration DB lookups into a single fetch (#652)
- Dreamer threshold and time-guard semantics:
check_and_schedule_dreamcount filter now includes onlydocuments.level == 'explicit'(dreamer-created levels are output, not input, and were inflating the threshold and creating a feedback loop);last_dream_atwrite relocated fromenqueue_dreamintoprocess_dreamso duplicate enqueues or failed runs no longer reset the 8-hour time guard (#573) - Deriver: blank observations are filtered out before embedding (previously triggered noisy embedding calls and persisted empty rows); blank-observation filtering unified across tool paths (#615)
- Surprisal module: filter for level observations changed from
{"level": levels}to{"level": {"in": levels}}—apply_filter()requires operator syntax, so the prior call silently returned 0 results and made the entire Surprisal phase of the Dream cycle a no-op (#581, fixes #559) - Removed hardcoded
stop_sequencesoverride from DeriverModelConfig(was clobbering operator-configured stop sequences) (#587) - Removed stale
stop_sequencesfrom tests (#607) - Embedding client:
embed()now wraps single-string input in an array, restoring compatibility with OpenAI-compatible third-party providers that reject scalar input (#586) - Docker Compose: deriver service startup gated on the API service healthcheck (prevents races where the deriver starts before the API has run migrations) (#689)
- Docker image:
HEALTHCHECKdirective removed from the shared base image — it probed an HTTP endpoint only the API serves, permanently marking deriver containers as unhealthy. Service-level health checks now belong in each service's own configuration (k8s readiness/liveness probes on the API Deployment only) (#530) tests/unified:--test-dir/--test-filearguments now use an argparse mutually-exclusive group instead of manual validation (#650)- CrewAI example updated for the latest CrewAI protocol (#631)
Removed
src/utils/clients.pydeleted; its responsibilities are split acrosssrc/llm/registry.py,src/llm/credentials.py, and the backend-specific modules (#459)HEALTHCHECKdirective removed from the shared Docker image (#530)
[3.0.6] - 2026-04-10
Changed
- Tightened transaction scopes across search, agent tools, queue manager, and webhook delivery to minimize DB connection hold time during external operations (#525)
- Search operations refactored to two-phase pattern — external work (embeddings, LLM calls) completes before opening a transaction (#525)
- Agent tool executor performs external operations before acquiring DB sessions (#525)
- Queue manager transaction scope reduced to only the critical section (#525)
- Webhook delivery no longer holds a DB session parameter (#525)
Fixed
- Session leakage in non-session-scoped dialectic chat calls (#526)
Added
- Health check endpoint (
/health) for container orchestration and load balancer probes (#510)
[3.0.5] - 2026-04-03
Fixed
- explicit rollback on all transactions to force connection closed
[3.0.4] - 2026-04-02
Added
- JSONB metadata validation enforces 100 key limit and max depth of 5 (#419)
Changed
- Schemas refactored from single
schemas.pyintoschemas/api.py,schemas/configuration.py, andschemas/internal.pywith backwards-compatible re-exports (#419)
Fixed
- Missing
deleted_atfilter onRepresentationManager._query_documents_recent()and._query_documents_most_derived()allowed soft-deleted documents to leak into the deriver's working representation (#456) CleanupStaleItemsCompletedEventemitted spuriously when no queue item was actually deleted (#454)- Empty JSON file uploads caused unhandled errors; now returns normalized error responses (#434)
- Memory leak:
_observation_locksswitched toWeakValueDictionaryto prevent unbounded growth (#419) - SQL injection in
dependencies.py: parameterizedset_configcalls to prevent injection via request context (#419) - NUL byte crashes: string inputs (message content, queries, peer cards) now stripped at schema level (#419)
- Filter recursion depth capped at 5 to prevent stack overflow (#419)
- Dedup-skipped observations now correctly reflected in created counts (#477)
- External vector store support for message search — routes queries through configured external vector store with oversampling and deduplication to handle chunked embeddings (#479)
- Dialectic agent no longer holds a DB connection during LLM calls — embeddings are pre-computed before tool execution, DB sessions isolated in
extract_preferences,query_documentsno longer accepts a DB session parameter (#477)
[3.0.3] - 2026-02-25
Added
- Consolidated session context into a single DB session with 40/60 token budget allocation between summary and messages
- Observation validation via
ObservationInputPydantic schema with partial-success support and batch embedding with per-observation fallback - Peer card hard cap of 40 facts with case-insensitive deduplication and whitespace normalization
- Safe integer coercion (
_safe_int) for all LLM tool inputs to handle non-integer values like"Infinity" - Embedding pre-computation and reuse across multiple search calls in dialectic and representation flows
- Peer existence validation in dialectic chat endpoints — raises ResourceNotFoundException instead of silently failing
- Logging filter to suppress noisy
GET /metricsaccess logs - Oolong long-context aggregation benchmark (synth and real variants, 1K–4M token context windows)
- MolecularBench fact quality evaluation (ambiguity, decontextuality, minimality scoring)
- CoverageBench information recall evaluation (gold fact extraction, coverage matching, QA verification)
- LoCoMo summary-as-context baseline evaluation
- Webhook delivery tests, dependency lifecycle tests, queue cleanup tests, summarizer fallback tests
- Parallel test execution via pytest-xdist with worker-specific databases
test_reasoning_levels.pyscript for LOCOM dataset testing across reasoning levels
Changed
- Workspace deletion is now async — returns 202 Accepted, validates no active sessions (409 Conflict), cascade-deletes in background
- Redis caching layer now stores plain-dict instead of ORM objects, with v2-prefixed keys, storage, resilient
safe_cache_set/safe_cache_deletehelpers, and deferred post-commit cache invalidation - All
get_or_create_*CRUD operations now use savepoints (db.begin_nested()) instead of commit/rollback for race condition prevention - Reconciler vector sync uses direct ORM mutation instead of batch parameterized UPDATE statements
- Summarizer enforces hard word limit in prompt and creates fallback text for empty summaries with
summary_tokens = 0 - Blocked Gemini responses (SAFETY, RECITATION, PROHIBITED_CONTENT, BLOCKLIST) now raise
LLMErrorto trigger retry/backup-provider logic - Gemini client explicitly sets
max_output_tokensfrommax_tokensparameter - All deriver and metrics collector logging replaced with structured
logging.getLogger(__name__)calls - Dreamer specialist prompts updated to enforce durable-facts-only peer cards with max 40 entries and deduplication
GetOrCreateResultchanged fromNamedTupletodataclasswithasync post_commit()method- FastAPI upgraded from 0.111.0 to 0.131.0; added pyarrow dependency
- Queue status filtering to only show user-facing tasks (representation, summary, dream); excludes internal infrastructure tasks
Fixed
- JWT timestamp bug —
JWTParams.twas evaluated once at class definition time instead of per-instance - Session cache invalidation on deletion was missing
get_peer_card()now properly propagatesResourceNotFoundExceptioninstead of swallowing itset_peer_card()ensures peer exists viaget_or_create_peers()before updating- Backup provider failover with proper tool input type safety
- Removed
setup_admin_jwt()from server startup - Sentry coroutine detection switched from
asyncio.iscoroutinefunctiontoinspect.iscoroutinefunction
Removed
explicit.pyandobex.pybenchmarks replaced by coverage.py and molecular.py- Claude Code review automation workflow (
.github/workflows/claude.yml) - Coverage reporting from default pytest configuration
[3.0.2] - 2026-01-27
Added
- Documentation for reasoning_level and Claude Code plugin
Changed
- Gave dreaming sub-agents better prompting around peer card creation, tweaked overall prompts
Fixed
- Added message-search fallback for memory search tool, necessary in fresh sessions
- Made FLUSH_ENABLED a config value
- Removed N+1 query in search_messages
[3.0.1] - 2026-01-27
Fixed
- Token counting in Explicit Agent Loop
- Backwards compatibility of queue items
[3.0.0] - 2026-01-19
Added
- Agentic Dreamer for intelligent memory consolidation using LLM agents
- Agentic Dialectic for query answering using LLM agents with tool use
- Reasoning levels configuration for dialectic (
minimal,low,medium,high,max) - Prometheus token tracking for deriver and dialectic operations
- n8n integration
- Cloud Events for auditable telemetry
- External Vector Store support for turbopuffer and lancedb with reconciliation flow
Changed
- API route renaming for consistency
- Dreamer and dialectic now respect peer card configuration settings
- Observations renamed to Conclusions across API and SDKs
- Deriver to buffer representation tasks to normalize workloads
- Local Representation tasks to create singular QueueItems
- getContext endpoint to use
search_queryrather than forcelast_user_message
Fixed
- Dream scheduling bugs
- Summary creation when start_message_id > end_message_id
- Cashews upgrade to prevent NoScriptError
- Memory leak in
accumulate_metriccall
Removed
- Peer card configuration from message configuration; peer cards no longer created/updated in deriver process
[2.5.1] - 2025-12-15
Fixed
- Backwards compatibility for
message_idsfield in documents to handle legacy tuple format
[2.5.0] - 2025-12-03
Added
- Message level configurations
- CRUD operations for observations
- Comprehensive test cases for harness
- Peer level get_context
- Set Peer Card Method
- Manual dreaming trigger endpoint
Changed
- Configurations to support more flags for fine-grained control of the deriver, peer cards, summaries, etc.
- Working Representations to support more fine-grained parameters
Fixed
- File uploads to match
MessageCreatestructure - Cache invalidation strategy
[2.4.3] - 2025-11-20
Added
- Redis caching to improve DB IO
- Backup LLM provider to avoid failures when a provider is down
Changed
- QueueItems to use standardized columns
- Improved Deduplication logic for Representation Tasks
- More finegrained metrics for representation, summary, and peer card tasks
- DB constraint to follow standard naming conventions
[2.4.2] - 2025-11-03
Fixed
- Langfuse tracing to have readable waterfalls
- Alembic Migrations to match models.py
- message_in_seq correctly included in webhook payload
Changed
- Alembic to always use a session pooler
- Statement timeout during alembic operations to 5 min
[2.4.1] - 2025-10-24
Added
- Alembic migration validation test suite
Fixed
- Alembic migrations to batch changes
- Batch message creation sequence number
Changed
- Logging infrastructure to remove noisy messages
- Sentry integration is centralized
[2.4.0] - 2025-10-09
Added
- Unified
Representationclass - vllm client support
- Periodic queue cleanup logic
- WIP Dreaming Feature
- LongMemEval to Test Bench
- Prometheus Client for better Metrics
- Performance metrics instrumentation
- Error reporting to deriver
- Workspace Delete Method
- Multi-db option in test harness
Changed
- Working Representations are Queried on the fly rather than cached in metadata
- EmbeddingStore to RepresentationFactory
- Summary Response Model to use public_id of message for cutoff
- Semantic across codebase to reference resources based on
observerandobserved - Prompts for Deriver & Dialectic to reference peer_id and add examples
Get Contextroute returns peer card and representation in addition to messages and summaries- Refactoring logger.info calls to logger.debug where applicable
Fixed
- Gemini client to use async methods
[2.3.3] — 2025-10-01
Changed
- Deriver Rollup Queue processes interleaved messages for more context
Fixed
- Dialectic Streaming to follow SSE conventions
- Sentry tracing in the deriver
[2.3.2] — 2025-09-25
Added
- Get peer cards endpoint (
GET /v2/peers/{peer_id}/card) for retrieving targeted peer context information
Changed
- Replaced Mirascope dependency with small client implementation for better control
- Optimized deriver performance by using joins on messages table instead of storing token count in queue payload
- Database scope optimization for various operations
- Batch representation task processing for ~10x speed improvement in practice
Fixed
- Separated clean and claim work units in queue manager to prevent race conditions
- Skip locked ActiveQueueSession rows on delete operations
- Langfuse SDK integration updates for compatibility
- Added configurable maximum message size to prevent token overflow in deriver
- Various minor bugfixes
[2.3.1] - 2025-09-18
Fixed
- Added max message count to deriver in order to not overflow token limits
[2.3.0] — 2025-08-14
Added
getSummariesendpoint to get all available summaries for a session directly- Peer Card feature to improve context for deriver and dialectic
Changed
- Session Peer limit to be based on observers instead, renamed config value to
SESSION_OBSERVERS_LIMIT Messagescan take a custom timestamp for thecreated_atfield, defaulting to the current timeget_contextendpoint returns detailedSummaryobject rather than just summary content- Working representations use a FIFO queue structure to maintain facts rather than a full rewrite
- Optimized deriver enqueue by prefetching message sequence numbers (eliminates N+1 queries)
Fixed
- Deriver uses
get_contextinternally to prevent context window limit errors - Embedding store will truncate context when querying documents to prevent embedding token limit errors
- Queue manager to schedule work based on available works rather than total number of workers
- Queue manager to use atomic db transactions rather than long lived transaction for the worker lifecycle
- Timestamp formats unified to ISO 8601 across the codebase
- Internal get_context method's cutoff value is exclusive now
[2.2.0] — 2025-08-07
Added
- Arbitrary filters now available on all search endpoints
- Search combines full-text and semantic using reciprocal rank fusion
- Webhook support (currently only supports queue_empty and test events, more to come)
- Small test harness and custom test format for evaluating Honcho output quality
- Added MCP server and documentation for it
Changed
- Search has 10 results by default, max 100 results
- Queue structure generalized to handle more event types
- Summarizer now exhaustive by default and tuned for performance
Fixed
- Resolve race condition for peers that leave a session while sending messages
- Added explicit rollback to solve integrity error in queue
- Re-introduced Sentry tracing to deriver
- Better integrity logic in get_or_create API methods
[2.1.2] — 2025-07-30
Fixed
- Summarizer module to ignore empty summaries and pass appropriate one to get_context
- Structured Outputs calls with OpenAI provider to pass strict=True to Pydantic Schema
[2.1.1] — 2025-07-23
Added
- Test harness for custom Honcho evaluations
- Better support for session and peer aware dialectic queries
- Langfuse settings
- Added recent history to dialectic prompt, dynamic based on new context window size setting
Fixed
- Summary queue logic
- Formatting of logs
- Filtering by session
- Peer targeting in queries
Changed
- Made query expansion in dialectic off by default
- Overhauled logging
- Refactor summarization for performance and code clarity
- Refactor queue payloads for clarity
[2.1.0] — 2025-07-17
Added
- File uploads
- Brand new "ROTE" deriver system
- Updated dialectic system
- Local working representations
- Better logging for deriver/dialectic
- Endpoint for deriver queue status
Fixed
- Document insertion
- Session-scoped and peer-targeted dialectic queries work now
Removed
- Peer-level messages
Changed
- Dialectic chat endpoint takes a single query
- Rearranged configuration values (LLM, Deriver, Dialectic, History->Summary)
[2.0.5] - 2025-07-11
Fixed
- Groq API client to use the Async library
[2.0.4] - 2025-07-02
Fixed
- Migration/provision scripts did not have correct database connection arguments, causing timeouts
[2.0.3] - 2025-07-01
Fixed
- Bug that causes runtime error when Sentry flags are enabled
[2.0.2] - 2025-06-27
Fixed
- Database initialization was misconfigured and led to provision_db script failing: switch to consistent working configuration with transaction pooler
[2.0.1] - 2025-06-26
Added
- Ergonomic SDKs for Python and TypeScript (uses Stainless underneath)
- Deriver Queue Status endpoint
- Complex arbitrary filters on workspace/session/peer/message
- Message embedding table for full semantic search
Changed
- Overhauled documentation
- BasedPyright typing for entire project
- Resource filtering expanded to include logical operators
Fixed
- Various bugs
- Use new config arrangement everywhere
- Remove hardcoded responses
[2.0.0] - 2025-06-24
Added
- Ability to get a peer's working representation
- Metadata to all data primitives (Workspaces, Peers, Sessions, Messages)
- Internal metadata to store Honcho's state no longer exposed in API
- Batch message operations and enhanced message querying with token and message count limits
- Search and summary functionalities scoped by workspace, peer, and session
- Session context retrieval with summaries and token allocation
- HNSW Index for Documents Table
- Centralized Configuration via Environment Variables or
config.tomlfile
Changed
- API route is now /v2/
- New architecture centered around the concept of a "peer" replaces the former "app"/"user"/"session" paradigm
- Workspaces replace "apps" as top-level namespace
- Peers replace "users"
- Sessions no longer nested beneath peers and no longer limited to a single user-assistant model. A session exists independently of any one peer and peers can be added to and removed from sessions.
- Dialectic API is now part of the Peer, not the Session
- Dialectic API now allows queries to be scoped to a session or "targeted" to a fellow peer
- Database schema migrated to adopt workspace/peer/session naming and structure
- Authentication and JWT scopes updated to workspace/peer/session hierarchy
- Queue processing now works on 'work units' instead of sessions
- Message token counting updated with tiktoken integration and fallback heuristic
- Queue and message processing updated to handle sender/target and task types for multi-peer scenarios
Fixed
- Improved error handling and validation for batch message operations and metadata
- Database Sessions to be more atomic to reduce idle in transaction time
Removed
- Metamessages removed in favor of metadata
- Collections and Documents no longer exposed in the API, solely internal
- Obsolete tests for apps, users, collections, documents, and metamessages
[1.1.0] - 2025-05-15
Added
- Normalize resources to remove joins and increase query performance
- Query tracing for debugging
Changed
/listendpoints to not require a request bodymetamessage_typetolabelwith backwards compatibility- Database Provisioning to rely on alembic
- Database Session Manager to explicitly rollback transactions before closing the connection
Fixed
- Alembic Migrations to include initial database migrations
- Sentry Middleware to not report Honcho Exceptions
[1.0.0] - 2025-04-10
Added
- JWT based API authentication
- Configurable logging
- Consolidated LLM Inference via
ModelClientclass - Dynamic logging configurable via environment variables
Changed
- Deriver & Dialectic API to use Hybrid Memory Architecture
- Metamessages are not strictly tied to a message
- Database provisioning is a separate script instead of happening on startup
- Consolidated
session/chatandsession/chat/streamendpoints
[0.0.16] - 2025-03-05
Added
- Detailed custom exceptions for better error handling
- CLAUDE.md for claude code
Changed
- Deriver to use a new cognitive architecture that only updates on user messages and updates user representation to apply more confidence scores to its known facts
- Dialectic API token cutoff from 150 tokens to 300
- Dialectic API uses Claude 3.7 Sonnet
- SQLAlchemy echo changed to false by default, can be enabled with SQL_DEBUG environment flag
Fixed
- Self-hosting documentation and README to mention
uvinstead ofpoetry
[0.0.15] - 2025-01-06
Added
- Alembic for handling database migrations
- Additional indexes for reading Messages and Metamessages
- Langfuse for prompt tracing
Changed
- API validation using Pydantic
Fixed
- Dialectic Streaming Endpoint properly sends text in
StreamingResponse - Deriver Queue handles graceful shutdown
[0.0.14] — 2024-11-14
Changed
- Query Documents endpoint is a POST request for better DX
Stringcolumns are nowTEXTcolumns to match postgres best practices- Docstrings to have better stainless generations
Fixed
- Dialectic API to use most recent user representation
- Prepared Statements Transient Error with
psycopg - Queue parallel worker scheduling
[0.0.13] — 2024-11-07
Added
- Ability to clone session for a user to achieve more loom-like behavior
[0.0.12] — 2024-10-21
Added
- GitHub Actions Testing
- Ability to disable derivations on a session using the
deriver_disabledflag in a session's metadata /v1/prefix to all routes- Environment variable to control deriver workers
Changed
- public_ids to use NanoID and internal ID to
use
BigInt - Dialectic Endpoint can take a list of queries
- Using
uvfor project management - User Representations stored in a metamessage rather than using reserved collection
- Base model for Dialectic API and Deriver is now Claude 3.5 Sonnet
- Paginated GET requests now POST requests for better developer UX
Removed
- Mirascope Dependency
- Slowapi Dependency
- Opentelemetry Dependencies and Setup
[0.0.11] — 2024-08-01
Added
session_idcolumn toQueueItemTableActiveQueueSessionTable to track, which sessions are being actively processed- Queue can process multiple sessions at once
Changed
- Sessions do not require a
location_id - Detailed printing using
rich
[0.0.10] — 2024-07-23
Added
- Test cases for Storage API
- Sentry tracing and profiling
- Additional Error handling
Changed
- Document API uses same embedding endpoint as deriver
- CRUD operations use one less database call by removing extra refresh
- Use database for timestampz rather than API
- Pydantic schemas to use modern syntax
Fixed
- Deriver queue resolution
[0.0.9] — 2024-05-16
Added
- Deriver to docker compose
- Postgres based Queue for background jobs
Changed
- Deriver to use a queue instead of supabase realtime
- Using mirascope instead of langchain
Removed
- Legacy SDKs in preference for stainless SDKs
[0.0.8] — 2024-05-09
Added
- Documentation to OpenAPI
- Bearer token auth to OpenAPI routes
- Get by ID routes for users and collections
- NodeJS SDK support
Changed
- Authentication Middleware now implemented using built-in FastAPI Security module
- Get by name routes for users and collections now include "name" in slug
- Python SDK moved to separate repository
Fixed
- Error reporting for methods with integrity errors due to unique key constraints
[0.0.7] — 2024-04-01
Added
- Authentication Middleware Interface
[0.0.6] — 2024-03-21
Added
- Full docker-compose for API and Database
Fixed
- API Response schema removed unnecessary fields
- OTEL logging to properly work with async database engine
fly.tomldefault settings for deriver setauto_stop=false
Changed
- Refactored API server into multiple route files
[0.0.5] — 2024-03-14
Added
- Metadata to all data primitives (Users, Sessions, Messages, etc.)
- Ability to filter paginated GET requests by JSON filter based on metadata
- Optional Sentry error monitoring
- Optional Opentelemetry logging
- Dialectic API to interact with honcho agent and get insights about users
- Automatic Fact Derivation Script for automatically generating simple memory
Changed
- API Server now uses async methods to make use of benefits of FastAPI
[0.0.4] — 2024-02-22
Added
- apps table with a relationship to the users table
- users table with a relationship to the collections and sessions tables
- Reverse Pagination support to get recent messages, sessions, etc. more easily
- Linting Rules
Changed
- Get sessions method returns all sessions including inactive
- using timestampz instead of timestamp
[0.0.3] — 2024-02-15
Added
- Collections table to reference a collection of embedding documents
- Documents table to hold vector embeddings for RAG workflows
- Local scripts for running a postgres database with pgvector installed
- OpenAI Dependency for embedding models
- PGvector dependency for vector db support
Changed
- session_data is now metadata
- session_data is a JSON field used python
dictfor compatibility
[0.0.2] — 2024-02-01
Added
- Pagination for requests via
fastapi_pagination - Metamessages
get_messageroutescreated_atfield added to each Table- Message size limits
Changed
- IDs are now UUIDs
- default rate limit now 100 requests per minute
Removed
- Removed messages from session response model
[0.0.1] — 2024-02-01
Added
- Rate limiting of 10 requests for minute
- Application level scoping