* feat(conclusions): expose reasoning level + allow filtering by level
The `level` of a conclusion (explicit / deductive / inductive /
contradiction) was filterable server-side but stripped from the
`Conclusion` response and not surfaced in either SDK. This adds it
end-to-end so callers can list explicit-only ("not dreamed on")
conclusions without dropping to raw HTTP.
- api: add `level` to the Conclusion response schema
- python sdk: `ConclusionLevel` type, `level` on Conclusion/response,
`level=` kwarg on ConclusionScope.list() and the async variant
- ts sdk: `ConclusionLevel` type, `level` on Conclusion/response,
`level` option on list()
- tests: assert level is exposed; add level-filter list test
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(conclusions): use generic filters= on list() instead of level= kwarg
Match the documented SDK convention (peers/sessions/messages all take a
generic `filters` dict passed through to the same dynamic server-side
filter logic) instead of a one-off `level=` kwarg. `level` filtering now
works as `list(filters={"level": "explicit"})` alongside any other
supported filter/operator.
The `level` field on the Conclusion response (added in the previous
commit) is kept — it's still not otherwise returned by the API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(conclusions): allow filtering by level on query() in py + ts SDKs
The branch's level-filter work exposed `filters=` on `list()` but left
`query()` (semantic search) hardcoding `{observer, observed}`, so callers
could filter the list endpoint by reasoning level but not semantic search —
asymmetric in both SDKs.
- Python: add keyword-only `filters` to `ConclusionScope.query` and
`ConclusionScopeAio.query`, merged over the scope's observer/observed.
- TypeScript: add optional `filters` arg to `ConclusionScope.query`,
mirroring the existing `list()` change.
The server `/conclusions/query` endpoint already honors filters in the body
(verified against production), so this is purely SDK surface parity.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(filters): document filtering conclusions by reasoning level
The using-filters page covered workspaces/peers/sessions/messages but not
conclusions. Add a "Filtering Conclusions" section showing level-based
filtering on both list() and query(), including the common "explicit only"
(exclude dream-derived) case and the in[deductive,inductive] inverse.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(conclusions): simplify filter merge to a single dict spread
Replace the merged_filters + if-block pattern in list()/query() (py sync,
aio, ts) with a single dict spread that layers the caller's filters over the
scope's observer/observed (and session). No behavior change — same merge
order (caller wins) — just less code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(conclusions): reject scope-managed keys in SDK conclusion filters
The generic filters= argument on ConclusionScope.list()/query() spread
user-supplied filters last, so a stray observer/observed/session key
silently overrode the scope and returned data from a different peer
pair. Add a fail-loud guard in both the Python and TypeScript SDKs that
rejects scope-managed filter keys with a clear error, directing callers
to peer.conclusions / conclusions_of(target) and the session= parameter.
session_id remains a valid filter on query() (which has no dedicated
session parameter).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: add exact content deduplication in document creation
* feat: add comment for index
* fix: harden times_derived logic across all callers to use max of inputs and existing + 1
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
The OpenAI backend passed tool_choice through raw while the Anthropic and
Gemini backends translate Honcho's canonical vocabulary to their native
form. On a mixed-provider fallback chain (e.g. Gemini primary -> OpenAI
backup), a canonical "any" reached OpenAI unchanged and was rejected as an
invalid param, since OpenAI only accepts none/auto/required.
Add a _convert_tool_choice to the OpenAI backend mirroring the others so a
single TOOL_CHOICE value resolves correctly regardless of which provider a
fallback lands on. "any"/"required" -> "required", auto/none pass through,
a tool-name string or {"name": ...} dict -> a function selection.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(llm): stop capturing live LLM clients in Langfuse generation spans
honcho_llm_call_inner is the @observe generation boundary, and default
auto-capture serialized every argument into the span input -- including
client_override (a live AsyncOpenAI/genai client) and selected_config
(which carries api_key). Auto-capture deep-copies the client into a
half-constructed object whose teardown raises:
- AsyncHttpxClientWrapper ... no attribute '_state' (OpenAI, stderr flood)
- BaseApiClient ... no attribute '_http_options' (Gemini, HONCHO-4HA)
and it leaked ModelConfig.api_key into traces.
Switch from auto-capture (denylist) to explicit annotation (allowlist):
disable capture_input/capture_output on the decorator and stamp curated,
serializable input (messages) and output (HonchoLLMCallResponse) via the
new annotate_current_generation_io helper. Full trace fidelity is
preserved; no client object or secret can reach a trace.
Fixes HONCHO-4HA
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(llm): track call tuning knobs as Langfuse model_parameters
Restore full trace fidelity after disabling @observe auto-capture: surface
every tuning knob (temperature, max_tokens, tools, reasoning effort, ...) on
the generation via model_parameters, sourced from the resolved effective
config instead of the raw function args.
Use a deny-list, not an allow-list: dump the whole ModelConfig and exclude
only secret-bearing fields (api_key, base_url, fallback, provider_params), so
new config knobs are traced automatically without keeping a hand-written list
in sync. The live client is never passed -- there is no useful trace
representation of it and serializing it is what triggered HONCHO-4HA.
Adds a deny-list test proving secrets never leak even when the config carries
a real api_key/base_url/provider_params (the production override-client path).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(llm): duplicate token usage to Langfuse + skip payload build when disabled
Mirror per-call token usage (input, output, prompt-cache read/creation) onto
the Langfuse generation via usage_details, so Langfuse renders native tokens
and cost in addition to the CloudEvents accounting.
Also guard both generation-annotation blocks behind settings.LANGFUSE_PUBLIC_KEY
so the model_dump-backed model_parameters payload (and the usage dict) are only
built when Langfuse is actually configured (addresses CodeRabbit: the annotate
helper no-ops when disabled, but the payload was still being constructed every
call).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: use git tags to fetch secrets for unified test
* fix: test failure
* fix: override AUTH_USE_AUTH and SENTRY_ENABLED
* fix: upload traces
* chore: rm run on PR
* fix: rm bucket from logs
* fix(agent_tools): strip display-format "id:" prefix from model-supplied observation IDs
Observations are presented to agents as [id:xxx], and models sometimes
copy the prefix verbatim despite tool-schema instructions to pass the
bare ID. This silently corrupts source_ids provenance on
create_observations_* (broken links stored in document metadata) and
breaks get_reasoning_chain lookups.
Normalize at both entry points. delete_observations is intentionally
not touched here since #746 already covers it.
Only the "id:" prefix is stripped: document IDs are nanoids whose
alphabet includes "-" and "_", so more aggressive cleanup could mangle
legitimate IDs.
Related to #719.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent_tools): strip whitespace remaining after "id:" prefix removal
Addresses CodeRabbit review: defends against "id: xxx" with a space
after the colon, and matches the docstring, which already promised
surrounding-whitespace stripping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: send OpenRouter app-attribution headers on OpenAI-compatible clients
Sets HTTP-Referer and X-Title on every AsyncOpenAI client constructed in
src/llm/registry.py (default, override-cached, and module-level CLIENTS) and
in the embedding client, so OpenRouter attributes Honcho's requests to the
"Honcho" app in its dashboard/analytics. Other OpenAI-compatible providers
ignore unrecognized headers, so this is safe to send unconditionally.
* fix: scope OpenRouter attribution headers to OpenRouter base URL only
Address review feedback on #805:
- Only inject attribution headers when the configured base_url starts
with https://openrouter.ai (via new _openrouter_headers() helper)
- Rename X-Title to X-Openrouter-Title per OpenRouter docs recommendation
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: drive default headers from base-URL map, drop embedding path
Replace the OpenRouter-specific _openrouter_headers helper with a generic
_DEFAULT_HEADERS_BY_BASE_URL prefix map + _default_headers_for lookup, so
OpenRouter always receives its attribution headers and another provider can be
added with a single map entry. Revert the embedding-client change (OpenRouter
has no embeddings endpoint, so that gate was dead code) and add a unit test for
the lookup helper.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: add model config option for json_object mode
* fix: catch possible validation error from structured output
* fix(llm): harden structured_output_mode json_object path
Follow-up fixes to the json_object structured-output mode for
OpenAI-compatible providers without Structured Outputs support:
- runtime: carry structured_output_mode onto the per-attempt fallback
config (select_model_config_for_attempt dropped it, silently sending
json_schema to a provider that can't parse it)
- backend: return a graceful empty on a contentless json_object
response instead of raising, matching the json_schema path, and
preserve token usage by normalizing the response
- backend: narrow the parse-failure catch to BadRequestError only, so
transient JSONDecodeError/ValidationError propagate to retry/fallback
instead of being swallowed to empty on the first attempt
- config: reject structured_output_mode on non-openai transports
(silent no-op otherwise); trim docs to the deriver, the only
structured-output feature
- backend: validate clean JSON before repair, cache the schema
instruction, and share json_object setup between complete()/stream()
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(llm): consolidate structured-output repair, drop dead seam
Fold the OpenAI backend's three structured-output repair sites
(LengthFinishReasonError, parsed=None, json_object) into the one shared
_parse_or_repair_structured_content helper, gated by an empty_on_missing
flag: json_object returns a graceful empty on a contentless response so a
loose provider can't crash the call, while json_schema raises so the
retry/fallback chain engages.
Delete the dead execute_structured_output_call seam and its only
collaborators (attempt_structured_output_repair, StructuredOutputFailurePolicy)
— it was never called and its single-shot validate/repair/empty model
conflicts with the retry behavior in honcho_llm_call.
No behavior change. Adds tests covering the json_schema parse fallbacks
(repair, refusal passthrough, no-content raise).
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* telemetry: use session and user IDs in langfuse
* test: update old span test
* fix: disable langfuse in unit tests
* fix: add post-loop synthesis span
* refactor: address PR review feedback on langfuse tracing
- Consolidate track_name onto LLMTelemetryContext as the sole home;
remove the honcho_llm_call kwarg and update 4 callers to set it on
telemetry directly. Sentry ai_track now reads telemetry.track_name.
- Decouple escaped-stream self-stamping from run-context exit ordering:
stream_final_response now resets _in_agent_run explicitly around drain.
- Narrow langfuse_agent_step wrap in the tool loop — between-turn
bookkeeping (iteration_callback, choice switch, increment) lifted
outside the span so it scopes only the LLM call + tools.
- Reword test conftest comment to behavior-only language.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: switch langfuse spans to imperative handles
Replaces the context-manager-based langfuse_agent_run/step with imperative
LangfuseAgentRun/Step handles so the run span can outlive the function that
opens it. Streaming responses now own the run handle from construction and
close it after drain, stamping the accumulated streamed text as trace output
(previously blank). Multi-turn generations always stamp provider/model and
step metadata, fixing the regression where only the first turn was annotated.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(llm): record effective prompt-only input on run span
The run-level Langfuse span recorded the raw messages parameter, which is
None for prompt-only calls. Mirror execute_tool_loop's handling and record
the synthesized user message so the trace input isn't blank.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(llm): drop StreamingResponseWithMetadata.__anext__ to prevent span leak
The standalone __anext__ delegated straight to the inner stream, bypassing
the token-folding and Langfuse run-handle close that live only in the
__aiter__ generator. Any caller driving the wrapper via anext() instead of
`async for` would leak the run span and lose final-stream token accounting.
Latent today (all callers use `async for`), removed to close the footgun.
Add tests covering the run-handle drain path: full drain stamps the
accumulated streamed text as the span output and closes once; an abandoned
stream still closes via the finally rather than leaking.
* chore(llm): document intentional empty-body propagate_attributes block
The `with propagate_attributes(...): pass` stamps the active @observe trace
root via the context manager's __enter__ side effect; the empty body reads
as deletable dead code. Add a comment so it isn't removed. Addresses PR review.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(llm): restore api.py types after __anext__ removal
Dropping StreamingResponseWithMetadata.__anext__ made it stop satisfying
the AsyncIterator protocol, breaking the result annotation and the
isinstance narrowing in honcho_llm_call. Widen the tool-less result
annotation to include StreamingResponseWithMetadata and narrow positively
to HonchoLLMCallResponse before reading .content.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* perf(reconciler): only trace Sentry transactions when work is found
The reconciler enqueues sync_vectors every ~5 min per deriver instance.
process_item wrapped every dequeued reconciler task in a single
process_reconciler_task transaction, so idle cycles (the common case,
where the cycle finds no rows and exits immediately) still created and
sampled a transaction + profile, draining Sentry tracing/profiling quota.
Remove the top-level transaction and push tracing into the sync batch
helpers, starting a per-batch transaction only after rows are confirmed.
Idle cycles now emit zero transactions; busy sweeps emit one smaller
transaction per batch operation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(telemetry): drop infra/scrape transactions via a Sentry traces sampler
Sentry was sampling every transaction at a flat traces_sample_rate with no
sampler. The Prometheus /metrics scrape endpoint alone accounted for ~92% of
all traced transactions (and their profiles), with /openapi.json and the
deriver metrics server adding more pure noise.
Add a traces_sampler that returns 0.0 for infra/scrape endpoints (/metrics,
/health, /openapi.json, /docs, /redoc, and metrics/openapi transaction names)
and the configured rate for real traffic. Sampling here (vs
before_send_transaction) means dropped transactions are never recorded or
profiled and the decision propagates to child spans. Shared init covers both
the API server and the deriver worker.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Peer- and session-scoped JWTs were effectively workspace-scoped: auth() walked the route's declared scope and fell through to a workspace match, so a {w: ws-a, p: alice} token could act on any peer in ws-a.
* feat: peer keys can read sessions they belong to; require workspace on scoped keys
* fix: authorize JWTs by narrowest scope and gate member reads
Follow-up hardening on the narrowest-claim auth fix:
- Scope get_peer_config member-read to the caller's own peer; a session
member could previously read a co-member's per-session config.
- Enforce session membership on POST /peers/{id}/chat: the session_id
arrives in the body (invisible to require_auth), so a peer key could
read any session's injected message history. Check is_peer_in_session
in the handler before the dialectic runs.
- Consolidate the workspace-match check in auth() to a single hoisted
guard so no branch can silently re-open cross-workspace access.
- Normalize empty-string scope claims to None in verify_jwt so a blank
workspace can't satisfy the peer/session token-shape invariant.
- Extract scope_requires_workspace(), shared by verify_jwt and the keys
API so the creation-time guard and verification invariant can't drift.
route requires auth) and CLAUDE.md auth-scoping guidance.
- docs: describe narrow-scope key semantics in the platform reference.
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix(dedup): reinforce times_derived on duplicate detection
times_derived was never incremented: the reject-new branch dropped the
reinforcement and the new-wins branch reset the count to 1, so the column
stayed pinned at 1 for nearly every conclusion. With every value equal,
ORDER BY times_derived DESC resolved to arbitrary heap order (oldest rows
first), which froze stale conclusions to the front of injected context.
- reject-new: increment existing_doc.times_derived
- new-wins: carry existing count forward onto the replacement
- add created_at DESC tiebreaker to both most_derived queries
* test(dedup): guard times_derived reinforcement + recency tiebreak
Three regression tests, each fails on pre-fix code:
- most-derived ties break toward recency, not insertion order
- rejecting a duplicate reinforces the surviving doc
- a winning duplicate inherits the replaced doc's count + 1
* fix(dedup): atomic reinforcement increment + deterministic tiebreak
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: defer embedding messages
* fix: rm gauges
* feat: embed messages immediately on create with reconciler fallback (#766)
Adds embed_messages_now background task so newly created messages are
searchable within seconds instead of waiting up to the reconciler
interval. Three-phase claim/lease → embed → persist never holds a DB
session across the embedding call; the reconciler remains the fallback
for failures and stragglers.
* fix: harden immediate-embed fast path and cover its error branches
Wrap embed_messages_now in a top-level try/except so a failure in the
claim or persist phase degrades to "reconciler will retry" instead of
escaping into the background-task runner; the rows stay pending+leased
and the reconciler heals them.
Add tests for the previously-uncovered branches: external-store-unavailable
persist path, the file-upload endpoint's embed scheduling, and direct unit
tests for the shared compute_chunk_positions / build_message_vector_record
helpers.
Document the semantic-search eventual-consistency window in search.mdx
(keyword matches are immediate; vector matches lag creation by seconds).
* fix: don't hold DB session across vector-store upserts
* fix: align semantic-search function to filter null rows
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: implement read DB and fix queue stale cleanup
* fix: use read_db in internal methods
* fix: mention read db in the CLAUDE.md
* fix: make TRACING checkout hook autocommit-safe; sample cleanup-gate jitter once
The DB.TRACING checkout hook ran `SELECT set_config(...)` at pool checkout,
before the dialect applies the read engine's AUTOCOMMIT isolation level. That
statement autobegins a transaction, and psycopg then refuses to switch the
connection into AUTOCOMMIT ("can't change 'autocommit' now: connection in
transaction status INTRANS"), so every read_only session 500s under TRACING and
the INTRANS connection leaks back to poison later write checkouts. Run the hook
in autocommit and restore the prior mode so it never leaves an open transaction;
set_config(..., is_local=false) is session-scoped and survives the boundary.
Add a regression test (fails without the fix) covering read_only + TRACING.
Also sample the stale-cleanup gate's jittered interval once per attempt instead
of re-rolling it every poll, so the spacing is a fixed deadline per cycle rather
than a random walk (and is testable at non-zero jitter ratios).
* fix: reset request_context in TRACING checkout-hook test
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: add generate_jwt.py script for creating scoped JWTs
Adds a CLI utility script for generating Honcho JWTs without needing
to call the /v1/keys API endpoint. Useful for local development and
bootstrapping admin tokens.
Features:
- --admin flag for full-access tokens
- --workspace / --peer / --session flags for scoped tokens
- --expires flag with human-friendly duration syntax (e.g. 5h, 30d, 1y)
- --print-only flag for scripting (outputs bare token)
Examples:
uv run python scripts/generate_jwt.py --admin
uv run python scripts/generate_jwt.py --admin --expires 24h
uv run python scripts/generate_jwt.py --workspace my-ws --expires 30d
uv run python scripts/generate_jwt.py --workspace my-ws --peer my-peer --expires 1y
* docs: document generate_jwt.py in README auth setup section
* fix: remove t='' override to preserve utc_now_iso default in JWTParams
Per CodeRabbit review: explicitly setting t="" bypasses JWTParams's
default utc_now_iso timestamp, causing tokens for the same scope to
become byte-identical. Omitting t lets the default apply, ensuring
each generated token is unique.
* fix: address JWT script review feedback
* fix: type, lint
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix(deriver): Remove connection retry logic and add jitter to polling interval
* chore(docs): Update changelog and document new configurations
* chore: increment version numbers
* feat(db): add connection retry, adaptive deriver polling, and pool metrics
Add resilience and visibility for DB connection handling under transaction-
pooler (Supavisor) saturation, where client-connection limits get exhausted
across many tenants.
- get_db/tracked_db now force an eager pool checkout with bounded exponential
backoff (tenacity), retrying SQLAlchemy TimeoutError + OperationalError so
transient pooler rejections degrade gracefully instead of 500ing. Toggle via
DB_CONNECTION_RETRY_ENABLED (+ delay/backoff knobs); ~10s default budget.
- Deriver polling backs off when idle or erroring (base -> max, x2 each cycle)
and snaps back to base on claimed work, cutting steady-state query load.
Toggle via DERIVER_POLLING_BACKOFF_ENABLED (+ max/multiplier).
- Add scrape-time db_pool_connections Prometheus gauge (checked_out/checked_in/
size/overflow, labeled api|deriver), registered in both the API lifespan and
the deriver metrics server.
- Make SqlalchemyIntegration explicit in both Sentry inits; wrap connection
acquisition in a db.pool.acquire span and capture live pool stats on
retry-exhaustion.
* feat(db): add acquisition counter and in-flight query gauge
Build on the pool-connection metrics with two signals that turn detection
into diagnosis under transaction-pooler saturation:
- db_connection_acquisitions{outcome=ok|retried|exhausted}: counts how often
connection checkout retries through pooler rejection — the alertable early
warning before requests start failing.
- db_queries_in_flight: statements actually executing on the wire (via
SQLAlchemy cursor-execute events, drift-proof across query errors). Pairs
with checked_out: the gap reveals connections held but parked (the "idle in
transaction during an external call" antipattern). Labeled namespace +
instance_type only; gated on METRICS.ENABLED for zero overhead when off.
Add DB-free unit tests for retry outcomes, polling backoff, and in-flight
gauge drift handling.
* fix: address CodeRabbit review on PR #758
- db: roll back the session on a retryable checkout failure before
retrying — a failed autobegin can leave it pending-rollback, making the
next db.connection() raise instead of re-checking-out cleanly. Cheap
Python-side cleanup when no connection was bound.
- metrics: guard DBPoolCollector.collect() so a pool-read/import hiccup
can't raise and abort the whole /metrics scrape (Prometheus drops ALL
metrics if any collector raises) — log and fall back to empty.
* fix(db): lazy retrying session + review fixes for connection backoff
Address Codex/CodeRabbit review on PR #758.
- Replace eager checkout with HonchoAsyncSession: a lazy AsyncSession that
checks out its connection (with retry) on the first DB-touching call, not at
construction. Request handlers doing non-DB work (embedding/file/LLM) before
their first query no longer pin a connection across it, while the API path
still gets checkout retry. Only the checkout is retried — the statement runs
once via super(), so writes are never duplicated. Tracing's set_config moves
into the same lazy acquire hook.
- Roll the session back on a retryable checkout failure before retrying, so a
failed autobegin can't leave it pending-rollback.
- Lower default POOL_TIMEOUT to 5s and validate it stays under the retry budget
for pooled (non-null) POOL_CLASS; update config.toml.example and v2/v3 docs.
- Clamp pool overflow gauge to >= 0 (was negative before the pool fills).
- Remove double-sleep in the deriver idle poll (true backoff cap, not 2x);
make in-flight instrumentation registration idempotent.
- Tests: HonchoAsyncSession lazy/idempotent acquire, statement-runs-once,
tracing, commit/rollback flag reset, get_db no-acquire-at-entry, polling-loop
single-sleep, and the POOL_TIMEOUT/retry-budget validator.
* fix(db): cover all DB-touching session methods; clear flag on close/reset
Address Codex follow-up review on PR #758 (polish, no behavior-critical bug).
- HonchoAsyncSession: wrap get/get_one/stream/stream_scalars/delete in addition
to execute/scalar/scalars/flush/merge/refresh/commit, so the "lazy checkout
with retry on first DB use" guarantee has no holes. connection() stays
unwrapped (acquire_connection_with_retry calls it — wrapping would recurse).
- Reset the acquired flag on close()/reset() too, so a session reused after
close/reset re-acquires (and re-wraps retry) on its next DB use.
- Fix stale comments: connection retry now applies lazily to the request path
via HonchoAsyncSession (config.py), and the FakeSession helper note.
- Tests: close/reset flag reset, and get/delete route through acquisition.
* feat: add new cloudevents for api routes
* fix: add total input tokens to RepresentationCompletedEvent
* feat(telemetry): inject honcho_version + emitter health metrics
* feat(telemetry): per-LLM-call event with try/finally emission + sampler
Adds LLMCallCompletedEvent (llm.call.completed) — fires once per provider hit
with full cost-attribution context: transport/provider_label, model, token
counts with cache breakdown, finish_reason, outcome (success or error),
is_final_attempt flag, retry/fallback state, duration, tool-call shape,
streaming flag, and agent correlation (run_id + iteration).
- src/telemetry/events/llm.py: new event class + CallPurpose closed enum
(deriver.representation, dialectic.answer, dream.deduction|induction,
summary.short|long). Resource id includes attempt so multi-attempt retries
in one iteration get distinct deterministic ids.
- src/telemetry/events/base.py: BaseEvent._volume_class ClassVar (default
"ground_truth"); the new event opts into "high_volume".
- src/config.py: TelemetrySettings.HIGH_VOLUME_SAMPLE_RATE (default 1.0).
- src/telemetry/emitter.py: deterministic sampler keyed on run_id (so an
entire agent trace is kept or dropped together). Aggregate envelopes
bypass the sampler. Sampled-out events increment the dedicated counter
separate from buffer_full/send_failed drops.
- src/llm/runtime.py: AttemptPlan gains attempt/retry_attempts/is_fallback
so the executor reads retry state without re-deriving it.
- src/llm/types.py: LLMTelemetryContext dataclass carrying workspace,
call_purpose, run_id, iteration, peer fields. Iteration is mutable so
the tool loop can set it per inner call.
- src/llm/executor.py: honcho_llm_call_inner wraps the backend call in
try/finally — emits on success AND on exception, with is_final_attempt
computed from AttemptPlan. Stream path emits a was_stream=True placeholder
(token totals deferred until streaming completion is wired through).
Telemetry failures swallowed.
- src/llm/api.py: threads telemetry kwarg through all 4 signatures into
both honcho_llm_call_inner and execute_tool_loop.
- src/llm/tool_loop.py: _telemetry_for_iteration helper copies the caller
context with iteration set per call — covers both the normal iteration
loop AND the max-iteration synthesis call (iteration N+1).
Tests cover success/error emission, sampler trace-coherence (same run_id →
same decision), volume_class enforcement, unknown call_purpose tolerance,
provider_label inference, and telemetry failure isolation. 378/378 pass.
* feat(telemetry): emit agent.iteration on every LLM response + synthesis
AgentIterationEvent was defined but never emitted on this branch. Phase 2
wires it up in execute_tool_loop so every LLM call inside an agentic loop
produces one event — including the no-tool terminating iteration and the
max-iteration synthesis call — and threads LLMTelemetryContext from dialectic
and dreamer specialists down through honcho_llm_call.
- src/telemetry/events/agent.py: AgentIterationEvent opts into
_volume_class="high_volume" so the Phase 1 sampler throttles it.
- src/llm/tool_loop.py: _emit_agent_iteration() helper fires once per
honcho_llm_call_inner response, BEFORE the no-tool early return so the
terminating iteration is counted. A second emission fires for the
max-iteration synthesis call BEFORE final_response is mutated with
cumulative totals (otherwise the per-iteration counts would double-count).
Emission is defensively skipped when telemetry context lacks run_id /
agent_type / parent_category / workspace_name; emit failures are swallowed.
- src/dreamer/specialists.py: BaseSpecialist.run passes LLMTelemetryContext
with parent_category="dream", agent_type=self.name, observer/observed,
call_purpose=f"dream.{self.name}".
- src/dialectic/core.py: _telemetry_context() builds a shared context for
both answer() and answer_stream(), using self._run_id (always set) +
workspace + observed peer.
Tests cover fresh-copy semantics, per-iteration vs terminating emission,
defensive skip cases, telemetry-failure isolation, and volume_class. 408/408
pass across telemetry + llm + utils + dreamer + dialectic.
* feat(telemetry): agent.tool.call.completed event + ToolResult metadata
Adds the missing generic per-tool-call event so read-only tools (search_*,
get_recent_history, get_observation_context, etc.) and the four existing
state-change tools all produce a telemetry record. Built on a new internal
ToolResult(content, metadata) contract so handlers can surface
search-specific fields (top_k/used_embedding/query_tokens/results_count)
to Phase 3 and create/delete counts to Phase 5's specialist rollups.
- src/telemetry/events/agent.py: AgentToolCallCompletedEvent at v1 with
_volume_class="high_volume". Resource id = {run_id}:{iteration}:{tool_call_seq}
so two calls to the same tool in one iteration don't collide
deterministic ids and get dedup-dropped downstream.
- src/utils/types.py: ToolResult dataclass; two new ContextVars
(_current_tool_call_seq + _last_tool_metadata) so tool_loop and the
execute_tool closure can communicate per-call telemetry without changing
the public Callable[[str, dict], Any] signature.
- src/utils/agent_tools.py: execute_tool times handlers, unwraps ToolResult,
publishes metadata, emits the event. Handlers updated to ToolResult
where useful: create/delete observations, update_peer_card, search_memory,
search_messages. Other handlers continue to return str.
- src/llm/tool_loop.py: set_current_tool_call_seq before each executor call;
read get_last_tool_metadata after and stash on all_tool_calls[i] for
Phase 5 rollups.
Tests cover ToolResult str-likeness, ContextVar round-trip, full-context
emission with search metadata, resource-id disambiguation, defensive skip
cases, telemetry isolation, truncation metadata, volume_class. 420/420 pass.
* feat(telemetry): RepresentationCompletedEvent v2 token breakdown + tool-less truncation
Bulks out the deriver's per-batch telemetry without bumping the event schema
version. New additive fields capture the full token breakdown (queued vs.
extra-context vs. scaffold), the cap configuration (batch_max_tokens,
max_input_tokens, was_flush_enabled), real cap-hit flags, and observer
fanout. `input_tokens` stays unchanged as the queued-message-tokens billing
key Xatu's Stripe meter reads.
The big enabler: src/llm/api.py now actually enforces max_input_tokens on
the tool-less LLM path. Before this, the deriver passed the kwarg but the
path silently dropped it — so the configured cap was advisory and
hit_input_token_cap couldn't be measured. Phase 4 wires truncation through
the same truncate_messages_to_fit helper the tool loop uses and surfaces
input_was_truncated on HonchoLLMCallResponse.
- src/telemetry/events/representation.py: 12 additive fields, schema_version
stays at 2.
- src/llm/types.py: input_was_truncated on HonchoLLMCallResponse.
- src/llm/api.py: tool-less path truncates messages before dispatch, flips
input_was_truncated on the response when clamping occurs. Split into
Literal[True]/Literal[False] branches for typecheck.
- src/deriver/queue_manager.py: QueueBatchResult dataclass replaces the
3-tuple return from get_queue_item_batch; carries hit_batch_token_cap
(computed from cumulative token sum vs cap), was_flush_enabled snapshot,
and batch_max_tokens. Worker loop unpacks + forwards.
- src/deriver/consumer.py: process_representation_batch gains the three
flag kwargs and forwards.
- src/deriver/deriver.py: derives the breakdown fields locally, populates
the new fields on emit, sources hit_input_token_cap from
response.input_was_truncated.
Tests cover schema stability, defaultable fields, input_tokens semantic
preservation, cap-hit flag round-trip, model_dump completeness, and
HonchoLLMCallResponse.input_was_truncated mutability. Existing
test_queue_processing.py tests updated for QueueBatchResult and mock
process_representation_batch signature. 479/479 pass.
* feat(telemetry): DreamRunEvent v2 scheduler reasons + DreamSpecialistEvent v2 rollups
Bumps both dream events to v2 with additive fields. DreamRunEvent gains
scheduler context (threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold / dream_type / enabled_types_count) threaded through the
dream queue payload — the two scheduler gates stay as separate fields rather
than collapsing into one trigger_reason, preserving the WHY-vs-WHEN
semantics. DreamSpecialistEvent gains denormalized rollups
(created_observation_count / deleted_observation_count / peer_card_updated /
search_tool_calls_count) sourced from Phase 3's ToolResult.metadata so the
counts reflect observation truth, not call truth.
- src/telemetry/events/dream.py: schema_version → 2 for both events; new
fields all defaultable so older producers still construct valid events.
- src/utils/queue_payload.py: DreamPayload + create_dream_payload accept
threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold.
- src/dreamer/dream_scheduler.py: check_and_schedule_dream computes the two
reasons at decision time and threads them through schedule_dream →
_delayed_dream → execute_dream → enqueue_dream.
- src/deriver/enqueue.py: create_dream_record / enqueue_dream gain the
kwargs and persist on the queue payload.
- src/dreamer/orchestrator.py: process_dream unpacks the payload; run_dream
accepts the kwargs and stamps them on DreamRunEvent.
- src/dreamer/specialists.py: BaseSpecialist.run walks response.tool_calls_made
and sums ToolResult.metadata.created_count / .deleted_count, sets
peer_card_updated, counts search-tool calls by name.
Tests cover schema_version bumps, defaultable Phase 5 fields,
threshold-vs-delay semantics, observation-vs-call-count rollup distinction,
and DreamPayload round-trip. Existing tests updated for the schema bump
and the new enqueue_dream kwargs. 488/488 pass.
* feat(telemetry): AgentToolSummaryCreatedEvent v2 token breakdown
Bumps schema_version to 2 and adds three additive breakdown fields so
analytics can answer "how much of a summary call's cost was the previous-
summary rollup vs. the new messages vs. the scaffold instructions".
- src/telemetry/events/agent.py: previous_summary_tokens, message_tokens,
prompt_scaffold_tokens added with sensible 0 defaults. input_tokens
retains its current semantic (provider-side LLM tokens) — the plan's
proposed `provider_input_tokens` was omitted because input_tokens
already serves that purpose and a duplicate would fork queries.
- src/utils/summarizer.py: emit now populates the three new fields from
values already in scope (messages_tokens, previous_summary_tokens,
prompt_tokens). Hoisted prompt_tokens calculation out of the
is_fallback conditional so both the save-summary path and the emit
share one binding — basedpyright couldn't prove the sibling-scope
binding was safe, and the compute is cheap + idempotent.
Tests cover schema bump, defaultable fields, input_tokens semantic
preservation, first-summary edge case, and breakdown round-trip.
493/493 pass.
* feat(telemetry): embedding.call.completed event + call-purpose ContextVar
Adds the final piece of cost-attribution telemetry: per-embedding-call
events covering every provider hit (single + batch + retry attempts).
Embedding calls are real provider spend that was invisible before this
phase; search-heavy paths (dialectic agentic) can produce more embedding
calls than LLM calls, so the new event participates in the shared
HIGH_VOLUME_SAMPLE_RATE.
- src/telemetry/events/llm.py: EmbeddingCallCompletedEvent at v1 with
_volume_class="high_volume". EmbeddingCallPurpose closed enum
(search_memory / search_messages / create_observations / vector_sync /
summary / message_create). Resource id = run:purpose:provider:model:input_count
so per-iteration calls in one agentic run don't collide.
- src/utils/types.py: _embedding_call_purpose ContextVar plus
@contextmanager wrapper. Nesting-safe via ContextVar.reset(token).
Callers wrap embedding-driving operations in
`with embedding_call_purpose("search_memory"): ...` — no changes to
the embedding client signature.
- src/embedding_client.py: _emit_embedding_call wraps each provider hit
with try/finally so success AND error paths emit. Errors propagate
unchanged. Each retry attempt of _process_batch emits its own event.
Unknown call_purpose slugs drop to None (validation against the enum
happens at emit time, not at context-manager-set time).
- src/utils/agent_tools.py: search_memory / search_messages /
search_messages_temporal / create_observations (batch + fallback) all
tag their embedding calls.
- src/crud/representation.py: save_representation tags with
CREATE_OBSERVATIONS; get_working_representation precompute tags with
SEARCH_MEMORY.
- src/crud/message.py: create_messages batch embed tags with
MESSAGE_CREATE; search_messages/temporal fallback tags with
SEARCH_MESSAGES.
Tests cover event shape, enum closure, ContextVar nesting/exception
cleanup, wrapper success+error emission, unknown-purpose graceful
fallback, telemetry-failure isolation. 550/550 pass across the full
telemetry+llm+utils+dreamer+dialectic+deriver+crud test set.
* chore: fix tests
* fix(telemetry): address review findings on stream events, context propagation, and cap detection
Five findings from a post-Phase-7 review (one resolved by the merge from
main, four addressed here):
- src/llm/executor.py: stream-path LLMCallCompletedEvent now fires AFTER
the stream is set up and drained (or on exception), with real duration
and accurate outcome. Previously the event was emitted before
execute_stream() ran and was always recorded as outcome="success" with
duration_ms=0, which silently masked stream-setup and stream-drain
failures. Wrapping the async generator in try/finally surfaces the real
outcome; token counts stay 0 because we still don't have them at stream
end (aggregate envelopes carry totals).
- src/deriver/deriver.py + src/utils/summarizer.py: deriver and summarizer
LLM calls now thread LLMTelemetryContext into honcho_llm_call. Before
this, the closed CallPurpose enum had DERIVER_REPRESENTATION /
SUMMARY_SHORT / SUMMARY_LONG slugs but those production call sites
didn't actually pass `telemetry=`, so their LLMCallCompletedEvents lost
workspace_name, parent_category, and call_purpose. summarizer threads
workspace_name through _create_and_save_summary → _create_summary →
create_short_summary / create_long_summary.
- src/utils/types.py + src/embedding_client.py: embedding_call_purpose
ctx manager now accepts workspace_name and run_id kwargs, backed by
two new ContextVars. EmbeddingCallCompletedEvent's publisher reads
both via get_embedding_workspace_name / get_embedding_run_id so
embedding events carry workspace and run correlation. All call sites
updated: search_memory / search_messages / search_messages_temporal /
_handle_create_observations_impl pass ctx.workspace_name +
ctx.run_id; create_observations standalone and create_messages pass
workspace_name; RepresentationManager.save_representation and
get_working_representation pass self.workspace_name.
- src/deriver/queue_manager.py: hit_batch_token_cap detection rewritten.
Previously summed kept-rows' token_count and checked against
batch_max_tokens, but the SQL filter `cumulative_token_count <= cap`
guarantees kept rows stay under the cap, so the flag almost never
fired. Now uses two follow-up queries: total token_count across the
included id range + EXISTS check for any session message past the
last-kept id. Both true → cap was actually binding.
(The fifth finding — deriver scaffold-token computation needing
estimate_deriver_prompt_tokens(custom_instructions) — was resolved by
the merge from main; the Phase 4 emit at src/deriver/deriver.py:283
already sources prompt_scaffold_tokens from the wrapped helper.)
567/567 telemetry+llm+utils+dreamer+dialectic+deriver+crud tests pass.
ruff + basedpyright clean.
* chore: ruff linting
* chore: clean AI generated comments references specs
* fix: address coderabbit changes
* fix(telemetry): address remaining PR review findings
Six findings from the PR 637 telemetry review batched into one commit.
- src/llm/executor.py + src/embedding_client.py: asyncio.CancelledError
now surfaces as outcome="cancelled" on both stream and sync paths,
distinct from "error". Client disconnects mid-stream and server
shutdowns are normal control flow and should not feed error-rate
alerting. LLMCallCompletedEvent and EmbeddingCallCompletedEvent
outcome Literal extended; docstrings + tests cover the new state.
- src/utils/types.py + src/llm/tool_loop.py: new iteration_scope()
context manager captures and resets the four per-tool-loop
ContextVars (_current_iteration, _current_tool_call_seq,
_current_provider_tool_call_id, _last_tool_metadata). Applied as a
typed decorator to execute_tool_loop so back-to-back loops in the
same asyncio Task (worker batches, tests) don't observe stale state.
- src/telemetry/events/api.py + src/routers/messages.py:
MessageCreatedEvent schema v1 → v2. Added required last_message_id
(nanoid public_id of the trailing message); get_resource_id now keys
on it instead of message_count, eliminating the collision case where
two same-size batches in the same session+source produced identical
event ids. message_count stays on the body for analytics.
- src/deriver/queue_manager.py: hit_batch_token_cap now computed from
the FINAL post-config-filter batch. Previously the flag used the
pre-filter messages_context[-1].id, which produced false positives
when _resolve_batch_configuration trimmed the trailing queue item —
telemetry reported a cap-hit when the actual returned batch was
short for unrelated reasons. Cap-detection block moved inside the
async with after the filter; no extra DB connection.
- src/config.py + src/telemetry/emitter.py: documented the
HIGH_VOLUME_SAMPLE_RATE orphan trade-off (rate<1.0 keeps aggregates
but drops children, so JOIN ON run_id queries see partial traces).
Behavior unchanged — rate defaults to 1.0.
- src/deriver/deriver.py: WARNING-level invariant logs when
response.input_tokens < messages_tokens (provider tokenization
drift) or prompt_scaffold_tokens <= 0 (estimator silent failure).
Best-effort — telemetry never bleeds into the deriver path
* fix(telemetry): stream retry, embed attempts, truncation, dedup
Address remaining audit findings on the cloudevents PR:
- Stream setup now runs inside the awaited honcho_llm_call_inner so
tenacity's retry wrapper in stream_final_response catches transient
setup failures (rate-limit, auth, network). Previously the returned
generator deferred execute_stream until first iteration — outside
the retry wrapper — crashing the request and bypassing telemetry.
- Embedding _emit_embedding_call gains an is_final_attempt parameter;
_process_batch threads the real retry index so dashboards stop
conflating one-shot, mid-retry, and exhausted-retry calls.
- _truncate_tool_output returns (text, original_chars, was_truncated)
and a new _maybe_truncated_result helper wraps in ToolResult when
truncation happens. Five handlers migrated. AgentToolCallCompletedEvent
fields was_truncated and result_chars_before_truncation are now
populated instead of always None/False.
- execute_tool_loop tracks any_iteration_truncated and stamps
input_was_truncated on the final response (both HonchoLLMCallResponse
and StreamingResponseWithMetadata). Dialectic now reports
hit_input_token_cap correctly.
- GetContextEvent.get_resource_id uses empty-string sentinel instead
of literal "none" so a peer named "none" can't collide with absent.
- generate_event_id folds honcho_version into the deterministic id so
same logical event from different deploys produces distinct ids.
* fix(telemetry): address audit findings across LLM/embed/event paths
Three rounds of telemetry audit findings, grouped by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). Previously the
inner generator deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation populated by _truncate_tool_output via
a new _maybe_truncated_result wrapper; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose context manager accepts parent_category.
- 4 new EmbeddingCallPurpose enum values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Queue batch cap detection (P2.1)
- hit_batch_token_cap keys on the pre-config-filter SQL boundary so the
"kept=900 of 1000 cap, next=300 excluded by cap" case reports True
while still avoiding the config-filter false positive.
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3) — top_k,
used_embedding, embedding_query_count, query_tokens, results_count.
Tests: stream-setup retry, stream-retry attempt sequence, post-stream
output_tokens write-back, is_final_attempt matrix, truncation E2E,
tool-loop hit_input_token_cap propagation, honcho_version in event id,
GetContextEvent disambiguation, queue_items_cleaned round-trip.
* fix(telemetry): address audit findings across LLM/embed/event paths
Four rounds of telemetry audit findings (initial + 3 follow-ups), grouped
by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). The inner
generator previously deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Queue batch cap detection
- hit_batch_token_cap previously required total_in_range >= cap, which
produced false negatives whenever the kept range didn't fully exhaust
the budget. Replaced with a pre-config-filter SQL boundary check
(P2.1), then further refined to a queue-item boundary comparison
(Fix 14) so trailing-context trimming doesn't false-negative either.
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation now populated by _truncate_tool_output
via _maybe_truncated_result; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11). except BaseException
catches cancellations too (Fix 16).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12), and uses ValidationException for unsupported types per
project guideline (Fix 17).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose accepts parent_category.
- 4 new EmbeddingCallPurpose values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Reconciler no longer holds DB session during embedding (Fix 15)
- _sync_documents and _sync_message_embeddings refactored into
three phases per CLAUDE.md guideline: fetch+detach in a small DB
scope, external embedding call without DB locks, writes in a fresh
short-lived DB scope. New _apply_*_sync helpers; orchestrators
expunge ORM objects before invoking. Vector store upsert + sync_state
updates stay in the apply phase together.
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3).
- Dialectic.prefetched_conclusion_count uses Representation.len() so
inductive + contradiction observations count too (Fix 10).
* fix(telemetry): orchestrator emit + review feedback
Three more rounds of audit findings + inline PR review, grouped:
Orchestration / emit reliability
- run_dream wrapped in try/finally so DreamRunEvent always emits, even
on unexpected exceptions including CancelledError (`finally` still
runs while cancellation propagates). Specialist except clauses
broadened from SpecialistExecutionError (never raised in src/) to
Exception so provider/DB/tool failures are recorded with
deduction_success=False / induction_success=False instead of crashing
past the emit.
- BaseSpecialist.run() telemetry state initialization + try/finally
hoisted above the preflight phase (peer lookup, peer-card preload,
create_tool_executor, get_model_config, prompt construction) so
preflight failures emit DreamSpecialistEvent(success=False) instead
of being dropped on the floor.
- Reverted the Round-4 _sync_documents / _sync_message_embeddings
phase split. The split introduced a race: rows were released from
FOR UPDATE SKIP LOCKED before the embed call, allowing two workers
to claim and clobber the same batch. Long-held DB transaction
restored (pre-existing CLAUDE.md violation accepted as a deliberate
trade-off; proper fix requires a claim/in_flight migration tracked
separately).
Schema + naming (PR-internal — none of these have shipped)
- threshold_reason → trigger_reason on DreamRunEvent, DreamPayload, and
every emit/scheduler/router/test call site (~45 src + 21 test lines).
Name now accurately reflects the field's role across "manual",
"surprisal", and "document_threshold" values.
- MessageCreatedEvent reset to schema v1 (was internally bumped to v2
for last_message_id but never shipped at v1 — downstream sees it
for the first time at merge).
- DreamSpecialistEvent gains created_counts_by_level /
deleted_counts_by_level: dict[str, int] keyed on the closed
level taxonomy. Per-tool-call events use list[str] (≤10 items),
but specialist runs aggregate 20+ — dict keeps emissions compact.
- QueueBatchResult marked frozen=True.
Per-call embedding attribution
- Agent tool embedding_call_purpose wraps for search_memory,
search_messages, search_messages_temporal, create_observations now
driven embedding cost rolls up under the right workflow.
- create_observations() signature gains parent_category kwarg
(mirrors existing run_id pattern).
Manual dream scheduling
- Manual /schedule_dream route now passes trigger_reason="manual" and
delay_reason="immediate". Previously both arrived as null in
DreamRunEvent, breaking analytics joins.
Queue-batch SQL perf
- next_exists_check folded into the main CTE query via
bool_or(cumulative_token_count > batch_max_tokens) OVER () in a
nested subquery. Cap detection is now one roundtrip per batch
instead of two.
Code/doc cleanup
- representation.py docstring uses generic "downstream metering key"
language (was "Xatu's Stripe meter"). bench runner --base-url help
uses a generic example host (was "groudon.fly.dev"). Public-facing
code/docs shouldn't reference internal service names.
Tests added for: orchestrator failure-path DreamRunEvent emission,
specialists preflight try/finally coverage, manual-dream
trigger_reason/delay_reason round-trip, dict-rollup accumulation across
multiple tool calls in a specialist run, CTE-fold one-roundtrip
behavior. Full Python suite passes (1236).
* fix(telemetry): correctness + attribution + emitter robustness
- Dreamer iteration count: read response.iterations directly so
one-shot runs no longer report iterations=0 and tool-using runs
include the terminal/synthesis LLM call.
- RepresentationCompletedEvent.observer_count counts successful
saves, not attempts.
- search_memory empty-memory fallback reports the snippet count when
message context is returned (was always 0).
- Wire parent_category through every embedding emit path: message
create (api), save_representation (representation), per-observation
fallback (caller-supplied), and the peer/session context routes
(api). get_working_representation accepts parent_category and
embedding_purpose so the internal fallback embed lands in the same
analytics bucket as the route-level precompute even when the
precompute is suppressed.
- BatchItem carries token_count so _process_batch reuses chunk-prep
counts instead of re-encoding every chunk for the telemetry proxy.
- Drop vestigial EmbeddingCallCompletedEvent.batch_size (always ==
input_count).
- Emitter: release the lock during HTTP send so a failing endpoint's
retry+backoff (~36s worst case) doesn't block other flushers;
edge-trigger the 80%-capacity warning so sustained backpressure
doesn't flood logs; defer event_id generation past the high-volume
sampler for events with run_id so sampled-out children don't pay
the sha256; harden emit() against sync callers with no running
loop; track threshold-flush tasks so shutdown() drains in-flight
sends before closing the HTTP client.
* fix(telemetry): tool cancellation emit, nanoid run_ids, version unification
- execute_tool: wrap post-work in finally so AgentToolCallCompletedEvent
fires on CancelledError; explicit handler sets is_error/result_str
before re-raising.
- run_id: replace str(uuid.uuid4())[:8] with generate_nanoid() across
dialectic/dreamer/specialists; matches project-wide nanoid convention.
- Bump _schema_version on events touched by run_id widening:
DialecticCompletedEvent v1→v2 (also covers hit_input_token_cap field),
AgentIterationEvent v1→v2, AgentToolConclusionsCreatedEvent v1→v2,
AgentToolConclusionsDeletedEvent v2→v3, AgentToolPeerCardUpdatedEvent
v1→v2.
- Unify honcho_version: single HONCHO_VERSION constant in src/_version.py
read from pyproject.toml (importlib.metadata fallback). Drop
TELEMETRY.HONCHO_VERSION setting. Use the constant for the FastAPI app
version (no more hardcoded "3.0.6") and for emitter body injection.
- Delete 17 tautological per-event test_schema_version methods; the
parametrized contract test still enforces version >= 1 across all events.
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat(embedding): add dimensions_mode for OpenAI dimensions= forwarding
Add EMBEDDING_MODEL_CONFIG__DIMENSIONS_MODE (auto|always|never) controlling
whether the dimensions= parameter is forwarded on OpenAI embeddings.create
calls. auto (default) sends it when the operator explicitly set
EMBEDDING_VECTOR_DIMENSIONS and the configured model is not on the
known-rejecting allowlist (currently text-embedding-ada-002).
The provenance check (was VECTOR_DIMENSIONS explicitly set?) lives as
EmbeddingSettings.resolve_send_dimensions() because it needs access to
model_fields_set, which the standalone resolver does not have. The
resolved boolean is passed into _EmbeddingClient at construction time;
the client never inspects mode or provenance.
Also pins cloudevents <2.0 — 2.0.0 reorganized the package and dropped
cloudevents.conversion and cloudevents.http, which src/telemetry/emitter.py
imports. The original `>=1.12.0` constraint allowed the broken 2.0 resolve.
With the pin, the imports resolve cleanly and the basedpyright warning
cascade (37+ warnings about unknown types) disappears.
Drive-by cleanups (all unnecessary cast/ignore comments flagged by
basedpyright after the cloudevents downgrade):
- vector_store/lancedb.py, tests/conftest.py, and
tests/deriver/test_vector_reconciliation.py — drop dead pyright ignores
- sdks/python/src/honcho/http/{async_,}client.py — drop unnecessary
cast(datetime, ...) (parsedate_to_datetime already returns datetime)
- vector_store/turbopuffer.py — cast(Any, rows) for the upsert_rows
TypedDict that the SDK exposes but our row builder doesn't satisfy
- tests/test_datetime_parsing.py — ignore reportArgumentType on the
test that deliberately passes wrong types to assert raises
* feat(models): honor EMBEDDING_VECTOR_DIMENSIONS in pgvector columns
* feat(startup): atomic swap dim-vs-MIGRATED guard for runtime schema validator
Add src/startup/embedding_validator.py that introspects the actual pgvector
column dim at boot and refuses to start if it does not match
EMBEDDING_VECTOR_DIMENSIONS. Runs after the DB pool is up and before the
embedding client is constructed, in both src/main.py (FastAPI lifespan) and
src/deriver/__main__.py.
Implementation details:
- Schema-qualified pg_attribute join through pg_class/pg_namespace respects
DB.SCHEMA rather than relying on search_path
- Bounded retry (3 attempts, 1s backoff) for transient introspection failure,
then fail-closed with "could not validate embedding schema" — uncertainty
is not a green light to serve traffic
- External-store sampler (turbopuffer, lancedb) enumerates workspaces from
the application DB and probes their lazy-created namespaces; current
per-namespace probe is a no-op stub since the SDKs do not expose
uniform dim introspection — full enumeration is left to
`configure_embeddings --report` in Phase 3
Atomic guard swap: deletes the old dim-vs-MIGRATED config validator (which
forbade non-1536 pgvector unless MIGRATED=True) in the same commit as the
new runtime validator. There is no release window where non-1536 pgvector
can start unprotected. The 9 dual-write branches that use VECTOR_STORE.MIGRATED
remain untouched and load-bearing for legacy-tenant backend swaps.
VECTOR_STORE_DIMENSIONS deprecation: drop the "must match" raise; in
propagate_namespace, check model_fields_set and emit logger.warning +
DeprecationWarning (DeprecationWarning alone is filtered by Python's default
config and would not reach operators). Always overwrite with
EMBEDDING.VECTOR_DIMENSIONS regardless.
Test changes:
- tests/test_models_vector_dim.py: Phase 1's VECTOR_STORE_TYPE=lancedb +
MIGRATED=true escape hatches removed; the test now passes on plain
EMBEDDING_VECTOR_DIMENSIONS=768
- tests/llm/test_model_config.py: the two tests asserting the old guards
replaced with tests for the new deprecation + acceptance behavior
- tests/startup/test_embedding_validator.py: 10 new tests — dim assertion
logic (pass/mismatch/missing/unbounded/non-public-schema), fail-closed
retry budget, real-test-DB pass, real-DB ALTER-then-validate, deprecation
warning capture, non-1536 + pgvector + MIGRATED=false at config time
* feat(scripts): add configure_embeddings bootstrap CLI
Adds scripts/configure_embeddings.py alongside the other one-off scripts
(provision_db, migrate_db, generate_jwt_secret, etc.). Invoked as
`uv run python scripts/configure_embeddings.py` — same convention as the
existing scripts in that directory, including the sys.path shim that
lets src.* imports resolve when run directly.
Bootstrap step for self-hosted installs at a non-default
EMBEDDING_VECTOR_DIMENSIONS — runs between `alembic upgrade head` and
starting the API/deriver.
pgvector ALTER safety (single transaction):
- LOCK TABLE {schema}.documents, {schema}.message_embeddings IN ACCESS
EXCLUSIVE MODE — closes the TOCTOU window between population check
and ALTER
- COUNT(*) WHERE embedding IS NOT NULL on both tables; refuse with a
non-zero exit if either is populated (ALTER ... USING NULL would
silently wipe those vectors)
- Snapshot HNSW index DDL from pg_indexes; drop, ALTER, recreate from
the captured DDL so operator-set HNSW params (m, ef_construction)
survive the round trip
External vector stores (turbopuffer, lancedb) are never created or
modified — namespaces are per-workspace and lazy-created on first write.
The --report mode enumerates workspaces and collections from the
application DB, derives the expected namespaces via
get_vector_namespace(), and prints a per-namespace status table.
CLI modes (mutually exclusive):
- (default) interactive: print plan, prompt to confirm
- --dry-run: print plan and exit 0 without touching the DB
- --yes: apply without prompt
- --report: print external-store namespace inventory and exit
Also updates src/startup/embedding_validator.py error-message paths and
docs/v3/contributing/configuration.mdx invocations to point at the new
script location.
Tests cover plan no-op, plan needs-alter, plan raises on missing column,
ALTER + HNSW round-trip, refuse-when-populated (monkeypatched count to
avoid wiring the full workspace/peer/collection/document FK chain just
to land one vector row), and idempotency.
* docs: add changing-embeddings operations page
Document the supported way to change EMBEDDING_VECTOR_DIMENSIONS or
EMBEDDING_MODEL_CONFIG__MODEL on a Honcho deployment: provision a new
deployment at the desired configuration, replay source data out of
band, cut over at the application layer.
The page explains the asymmetry:
- Dimension is machine-enforced as immutable. The startup validator
introspects pg_attribute and crashes the API/deriver on mismatch.
- Model is operator-owned. There is no persistent metadata recording
which model produced each vector, so a same-dim model swap is
silently undetectable — flagged with a Warning callout.
Also documents the truncation edge case (text-embedding-3-large truncated to 1536 with EMBEDDING_VECTOR_DIMENSIONS left at default)
and the DIMENSIONS_MODE=always mitigation, plus a pointer that
storage-backend swap (VECTOR_STORE_MIGRATED + reconciler) is a distinct operation unaffected by this work.
Registers the page in docs/docs.json under the Self-Hosting nav group
and cross-links from configuration.mdx.
* fix(embedding): correct turbopuffer regex + tighten DIMENSIONS_MODE docs
- Turbopuffer attribute type for a vector column is `[N]f32` / `[N]f16` /
`[N]i8`, not `f32_vector(N)` as the earlier probe assumed. The earlier
regex returned None for the real SDK format, so existing Turbopuffer
namespaces would have been reported as "missing" instead of validated
for mismatch. Regex switched to `\[(\d+)\]` which is the
vendor-stable shape. Test cases rewritten to lock the actual format.
- docs/v3/contributing/configuration.mdx had a contradictory pair of
bullets: 223 said explicit 1536 makes `auto` forward dimensions=, 224
said `auto` would skip the parameter because 1536 is the default.
Operators reading both would (rightly) conclude they need `always`
even when `auto` would work. Rewrote both bullets so:
- `auto` is provenance-driven (explicit-set, not non-default-value).
- `always` is positioned as defense-in-depth for config layers that
might strip explicit default-valued envs, not the only path for
same-as-default truncation.
* fix(embedding): address PR #678 review comments
CodeRabbit + Rajat review feedback. All actionable items addressed
except two false-positives (responded on PR).
Bug fixes:
- deriver telemetry leak: validator was called outside try/finally so
shutdown_telemetry() did not run on validation failure. Moved inside.
- _emit_report printed "no effect with pgvector" unconditionally,
including from implicit post-apply calls. Added is_report_mode flag;
only print on explicit --report.
- LanceDB and Turbopuffer probes returned None when the namespace
existed but its schema was malformed (no vector field / unparseable
type string), silently bucketing real corruption as "missing"
(lazy-create) and letting it pass the startup validator. Now raise
VectorStoreError with actionable diagnostics; None remains valid only
for "namespace does not exist."
- Startup validator only sampled message namespaces; added a parallel
Collection-row sample so document namespaces are probed too, with the
same dim assertion. Mirrors the --report path.
Hygiene:
- StartupValidationError now subclasses HonchoException so existing
exception handlers recognize it. ValidationException is @final and
has 422 request-validation semantics that would be misleading here.
- scripts/configure_embeddings.py main() no longer spins up two event
loops. engine.dispose() moved into a try/finally inside _async_main
so cleanup runs in the same loop as the pipeline.
- Replaced hand-rolled retry loop with tenacity.AsyncRetrying; same
fail-closed semantics, less code, before_sleep_log for visibility.
- Added _validate_identifier() defense-in-depth: DB.SCHEMA and HNSW
index names are regex-checked against [A-Za-z_][A-Za-z0-9_]* before
SQL interpolation. Operator config + DB catalog are not user input
under the current threat model, but the constraint is cheap to gate.
Test + docs:
- test_app_settings_accepts_non_1536_with_any_vector_store_configuration
now actually exercises turbopuffer (was missing); supplies a dummy
TURBOPUFFER_API_KEY to satisfy the model_validator.
- changing-embeddings.mdx: hyphenated "out-of-band" per reviewer style.
* fix: modify conftest to fix ci
* fix: ci tests for typescript server
The previous mutual-exclusion check compared --test-dir against its
default string literal, so passing --test-file together with an
explicit --test-dir tests/unified/test_cases silently bypassed the
check. Replace with argparse.add_mutually_exclusive_group() and apply
the default path post-parse so the bare invocation still works.
* docs(integrations): add @honcho-ai/vercel-ai-sdk guide
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(integrations): rewrite Vercel AI SDK guide as cookbook style (DEV-1485)
Reshapes the guide to cookbook formula, adds Full Script section, fixes
maxSteps → stopWhen for ai-sdk v5, renames package, and prunes stale notes.
See PR for full decision log.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(integrations): lead Vercel AI SDK verification with direct-inspection check
- Restructure Verifying section: direct inspection (token delta + dashboard) is now step 1 so readers isolate Honcho's contribution before grading model behavior
- Behavioral tests (first turn, multi-turn, cross-session, tool calling) follow as steps 2-5
- Note `result.toolCalls` as the way to confirm which Honcho tool fired (tool names don't appear in `result.text`)
- Signpost the Full Script from Complete Example so the two snippets read as a staircase, not a duplicate
Addresses review comments on PR #635.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(tests): satisfy basedpyright in test_representation_manager
The save-representation tests added in #615 were structurally correct but
failed strict typing in two places. Static Analysis has been red on main
since the merge.
- `mock_save.await_args` is `_Call | None`; assert it's not None before
reading `.kwargs` / `.args` so basedpyright can narrow the type
- `SimpleNamespace(...)` passed as `message_level_configuration` is an
intentional duck-typed mock (only `.dream.enabled` is read by
`save_representation`), so opt out at the call site with
`# pyright: ignore[reportArgumentType]` rather than constructing a
full `ResolvedConfiguration` (matches the existing `reportPrivateUsage`
ignore pattern in this file)
No runtime behavior changes; `uv run basedpyright` is now clean
project-wide.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(tests): pad timestamp windows in test_messages for clock skew
Three timestamp tests captured `before_request` / `after_request` with
`datetime.now(UTC)` on the host and asserted the server's `created_at`
fell within. Under Docker, the Postgres container's clock can skew tens
of ms from the macOS host, flipping the assertion intermittently under
parallel pytest load.
Pad each window by 1 second on both sides — wide enough to absorb
realistic skew, narrow enough that the test still proves the timestamp
is server-current.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(integrations): tighten Verifying section after end-to-end smoke
Smoke-tested all five verification steps against a fresh Sonnet 4.6 + Honcho integration. Three findings, all reflected here:
- Cross-session recall (#4): added Note about DERIVER_REPRESENTATION_BATCH_MAX_TOKENS=1024 — short warmups don't accumulate enough content to flush observations, so cross-session recall returns empty even on a working integration.
- Tool calling prompt (#5): replaced the honcho_chat patterns prompt with a verbatim-retrieval honcho_search prompt. Sonnet skips honcho_chat when middleware-injected context already answers; verbatim retrieval forces a fire.
- Tool inspection (#5): replaced result.toolCalls reference with result.steps[i].toolCalls + flatMap snippet. Top-level toolCalls is empty in multi-step calls (stopWhen: stepCountIs(N)) — the fires are nested inside steps.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(integrations): make Step 4 cross-session test durable via honcho_search
Replace the prose-recall test ("Based on what we've talked about, what do you know about me?") with a forced honcho_search call. Prose recall depended on the model getting deriver-built representation/peer-card in its system prompt, which is gated behind DERIVER_REPRESENTATION_BATCH_MAX_TOKENS=1024 — short tutorial-length conversations don't trigger it, producing false negatives on a working integration.
honcho_search hits message embeddings, which are computed synchronously at message persist time (src/crud/message.py:262-276), so peer-scoped retrieval works regardless of how short the prior session was. Also folds the result.steps[i].toolCalls inspection snippet from the old Step 5 into Step 4 — same prompt, no need for two sections.
Drops Step 5 entirely.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix(dreamer): threshold and time-guard semantics
Finding 2: filter count_stmt on documents.level == 'explicit' in
check_and_schedule_dream. Dreamer-created levels (deductive, inductive,
contradiction) are consolidation output, not input, and would otherwise
inflate the threshold count and create a feedback loop.
Finding 3 (code-level): relocate last_dream_at write from enqueue_dream
(enqueue.py) to process_dream (orchestrator.py), inside the
'if result is not None' block. Duplicate enqueues can no longer reset
the 8-hour time guard clock. Failed/never-run dreams don't advance it.
Success criteria: lenient (any non-null DreamResult counts). Pending
Vineeth confirmation — will adjust to strict/middle if requested.
Tests pending in follow-up commits.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(dreamer): threshold filter + last_dream_at relocation regression tests
Tests for Finding 2 and Finding 3 (code-level):
- TestThresholdFilter (tests/dreamer/test_dream_scheduler.py):
* Mixed levels below explicit threshold: 30 explicit + 40 deductive
+ 10 inductive → no trigger (core regression, buggy count would trigger)
* Explicit-only at threshold: 60 explicit → triggers
* Contradiction excluded: 100 contradiction + 10 explicit → no trigger
(confirms positive == "explicit" filter excludes all dreamer output)
- TestEnqueueDreamMetadataShape (tests/deriver/test_enqueue_dream.py):
* AsyncMock-patched update_collection_internal_metadata verifies
enqueue writes last_dream_document_count but NOT last_dream_at
- TestLastDreamAtCompletionWrite (tests/dreamer/test_dreamer_integration.py):
* Happy path: run_dream returns DreamResult → last_dream_at written
* Failure path: run_dream returns None → last_dream_at absent
* Exception path: run_dream raises → last_dream_at absent,
process_dream swallows exception (queue-processed semantics preserved)
Docstring on check_and_schedule_dream tightened: "document threshold"
-> "explicit-observation threshold" to reflect filter semantics.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dreamer): preserve last_dream_document_count in completion write
CodeRabbit caught this: update_collection_internal_metadata uses a
top-level JSONB `||` merge, so passing {"dream": {"last_dream_at": ...}}
replaces the entire "dream" subkey and drops last_dream_document_count
that was written by enqueue_dream.
Symptom: after every completed dream, the baseline drops to 0. Next
check_and_schedule_dream reads documents_since_last_dream as
current_count - 0 = current_count, so any collection with >= 50
explicit observations can re-trigger immediately once the 8h guard
expires, even with no new raw material.
Fix: read-modify-write. Fetch current collection, merge last_dream_at
into the existing "dream" dict, write the merged dict back. Preserves
sibling keys (current: last_dream_document_count; future-proof for
telemetry fields that might land in PR 4).
Regression test added to tests/dreamer/test_dreamer_integration.py:
pre-seeds {"dream": {"last_dream_document_count": 42}}, runs
process_dream, asserts both last_dream_at is written AND
last_dream_document_count == 42 is preserved.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dreamer): address CodeRabbit feedback on b89997c
- enqueue.py: read-modify-write preserves last_dream_at when writing baseline
- dream_scheduler.py: explicit-level filter on execute_dream count query
- test fixture: pin DOCUMENT_THRESHOLD and ENABLED_TYPES for stability
- integration test: timezone-aware assertion on last_dream_at
Regression test added for enqueue sibling-drop (symmetric to c8fe40a).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dreamer): session lookup symmetry + row lock on dream metadata RMW
- dream_scheduler.py: explicit-level filter on execute_dream session lookup
(baseline and session pick must agree on the same document set)
- crud.collection.get_collection: optional with_for_update flag for callers
that need serialized read-modify-write on internal_metadata
- enqueue.py + orchestrator.py: pass with_for_update=True on the RMW reads
to close the TOCTOU between concurrent enqueue and completion writes
Follow-up filed for jsonb_set-based nested updates (docs/factory/backlog/).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dreamer): explicit-only count on manual schedule_dream route
The third caller of enqueue_dream — POST /workspaces/{id}/schedule_dream —
was passing an all-levels document count as the baseline, breaking symmetry
with check_and_schedule_dream and execute_dream after Loop 2's filter fixes.
Filter the manual route's count to match.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(dreamer): document explicit-only invariant on enqueue_dream.document_count
Loop 3 follow-up on d76627a. The parameter's semantic tightened across Loop
2 (check_and_schedule_dream, execute_dream) and Loop 3 (schedule_dream route)
to "explicit-level count, used as the baseline," but the signature still read
"Current document count for metadata update." The next caller would have no
way to know from the function contract.
Docstring now spells out: (1) the value is explicit-only, (2) it's written
as last_dream_document_count, (3) it's the baseline that
check_and_schedule_dream subtracts from to compute
documents_since_last_dream, (4) passing a count that includes non-explicit
levels (deductive, inductive, contradiction) inflates the baseline and
suppresses the next scheduled dream.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(dreamer): rename current_document_count → current_explicit_count
Loop 3 follow-up on d4e10e3. After Loop 2's filter landed, the local in
check_and_schedule_dream held an explicit-only count but was still named
current_document_count — asymmetric with execute_dream's current_explicit_count
(line 201) and contradicting the filter on line 269 that produces the value.
Pure rename: three occurrences (definition at 271, subtraction at 274, log
extra key at 282). No test references. Naming-as-invariant alignment with
d76627a (query filters), d4e10e3 (parameter docstring), and Loop 1's local
rename in execute_dream.
The persisted JSONB key last_dream_document_count is the one remaining
drift-layer; filed as plastic-claudebook backlog item for a separate PR
with an intentional migration path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dreamer): atomic guard-pair write + in-flight stampede defense
Loop 4 response to Vineeth's CHANGES_REQUESTED on PR #573.
The pre-Loop-4 enqueue-time write of last_dream_document_count was serving
double duty: rate limiter AND stampede latch. By arming the 8h guard the
moment a dream entered the pipeline, it implicitly blocked a second dream
from being scheduled during the in-flight window. Loop 3 relocated the
last_dream_at write to completion without moving its sibling baseline,
splitting the semantic pair and exposing the latch role that had lived
only in Vineeth's head.
Invariant (now pinned to check_and_schedule_dream's docstring): from the
moment a dream is scheduled until it completes or fails, no second dream
may be enqueued for the same (workspace, observer, observed) — and the
baseline count advances only when consolidation actually happened.
Changes:
- enqueue_dream: remove the last_dream_document_count write entirely and
drop the document_count parameter. enqueue no longer touches dream
metadata; the implicit stampede latch is replaced by an explicit
queue-backed defense.
- process_dream: extend the existing row-locked RMW to write both guard
fields atomically. Current explicit-doc count is recomputed inside the
locked block (not carried on DreamPayload) so the pair reflects the
actual consolidation moment.
- check_and_schedule_dream: query QueueItem for pending dreams on this
collection's work_unit_keys (mirrors uq_queue_dream_pending_work_unit_key)
before arming a timer. Uses queue state as source of truth rather than
reflecting it into metadata.
- Tests: two new coherence tests under TestGuardPairCoherence —
test_pending_queue_item_blocks_second_schedule walks the stampede timeline,
test_silent_failure_allows_retry_on_same_corpus verifies failed dreams
don't consume the baseline. Existing tests updated to the new contract.
* chore(dreamer): trim comment slop from loop-4 atomic pair work
Compress three verbose comments added in d24958d — the invariant itself
is captured in check_and_schedule_dream's docstring, so the inline
narrative restates what the code already says.
- dream_scheduler.py defense C block: 5 lines → 2
- orchestrator.py atomic pair write: 4 lines → 1
- enqueue.py docstring paragraph: 5 lines → 2
Net: +5/-14. Follows Eri's eef27be precedent on sillytavern-honcho PR #7.
---------
Co-authored-by: lilyplasticlabs <lily@plasticlabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(deriver): ignore blank observations before embedding
* Address PR review on observation normalization
* Harden mock await arg access in tests
* Unify blank observation filtering across tool paths
* Move soft-delete query test back to fixture class
* fix: Add JSON repair for truncated LLM responses across all providers and Gemini thinking budget support
LengthFinishReasonError from OpenAI-compatible providers (custom, openai, groq) was crashing the deriver
with 14k+ occurrences in production. The vLLM path already had repair logic but it was gated on
provider=="vllm", unreachable when routing through litellm as a custom provider.
- Extract shared _repair_response_model_json() helper for all providers
- Catch LengthFinishReasonError in OpenAI/custom parse() path and repair truncated JSON
- Add repair fallback to Anthropic and Gemini response_model paths
- Add repair fallback to Groq response_model path
- Pass thinking_budget_tokens to Gemini 2.5 models via thinking_config
- Add 14 tests covering repair paths for all providers and Gemini thinking budget
Fixes HONCHO-YC
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: live llm integration tests
* feat: Consistent Model Config Protocol
* fix: migrate the remaining app callers off the legacy llm_settings path
* fix: Docs and regression tests
* fix: refactor llm runtime path to model-config-only API
* fix: refactor config to nested model-config source of truth
* fix: refactor llm streaming and tool dispatch through backends
* fix: cut over llm config to nested model_config only
* fix: collapse vllm and custom into openai_compatible transport
* feat: refactor llm config to explicit transports and bare model ids
* feat: (embed) Add configurability for embedding model
* fix: tests for embedding provider
* fix: Address Review Comments
* fix: (llm) remove Groq backend and per-vendor base URLs
* chore: move llm tests
* fix: (llm) address review findings — config regressions, backend bugs, dead code
* fix: address backend end silly errors
* chore: (docs) update configuration and self-hosting guides
* chore: fix tests
* fix: address code rabbit comments
* fix: add validation to the dream settings
* fix: further address code rabbit comments
* fix: Address Code Rabbit Comments
* fix: Another round of code rabbit
* fix: Address Code Rabbit Nits
* fix: tests
* refactor: rename thinking validator to reflect transport scope
_validate_anthropic_thinking_minimum only enforces the >=1024 rule for
Anthropic and no-ops for other transports, so the name was misleading
now that it's shared across ConfiguredModelSettings, FallbackModelSettings,
and ModelConfig. Renamed to _validate_thinking_constraints with a docstring
clarifying per-transport behavior. No logic change.
* fix(config): drop transport-specific thinking params when env override changes transport
_fill_defaults_for_nested_field previously preserved the default MODEL_CONFIG's
thinking_budget_tokens/thinking_effort across a transport override. This leaked
Gemini-family defaults (e.g. thinking_budget_tokens=1024) into OpenAI-transport
overrides, and the OpenAI backend then correctly rejected the unsupported param
at call time (OpenAI uses reasoning.effort, not a token budget).
The helper now strips thinking_budget_tokens and thinking_effort from the
default dict when the env override supplies a transport different from the
default's. Explicit thinking params in the override are preserved.
* fix(config): apply thinking-param strip to dialectic level merge too
DialecticSettings._merge_level_defaults does its own inline MODEL_CONFIG
merge (parallel to _fill_defaults_for_nested_field), so the previous fix
missed dialectic-level overrides. E.g. flipping
DIALECTIC_LEVELS__minimal__MODEL_CONFIG__TRANSPORT from gemini (default)
to openai still leaked the default thinking_budget_tokens=0 into the
openai config, which the OpenAI backend then rejected at call time.
The level-merge path now applies the same 'strip transport-specific
thinking params when transport changes' rule as the generic helper.
Added a regression test exercising the merge validator directly.
* refactor(llm): wire ModelConfig knobs through, prune clients.py migration leftovers
Three connected fixes to finish carving the LLM stack out of src/utils/clients.py
and into src/llm/:
1. Propagate ModelConfig tuning knobs into backend calls.
honcho_llm_call_inner built extra_params from only {json_mode, verbosity},
silently dropping top_p, top_k, frequency_penalty, presence_penalty, seed,
and operator-supplied provider_params from any ModelConfig. Thread the
selected config through ProviderSelection and merge
build_config_extra_params(selected_config) into extra_params; per-call
kwargs still win over provider_params defaults. Makes
_build_config_extra_params public as build_config_extra_params so
clients.py and request_builder.py share one translation. Adds
TestModelConfigExtraParamsPropagation covering OpenAI/Anthropic knob
propagation, provider_params passthrough, and per-call override
precedence.
2. Drop dead extract_openai_* duplicates in clients.py.
extract_openai_reasoning_content, extract_openai_reasoning_details, and
extract_openai_cache_tokens had no callers outside their own definitions
— the live implementations live in src/llm/backends/openai.py. -103
lines from clients.py.
3. Unify on ModelTransport, delete SupportedProviders.
The "google" vs "gemini" split forced a _provider_for_model_config
translation shim in two places. Replace all SupportedProviders usages
with ModelTransport, rename CLIENTS["google"] → CLIENTS["gemini"],
update provider branches + LLMError labels + reasoning-trace entries
accordingly. Trace JSONL now writes "provider": "gemini" instead of
"google" — consistent with the broader env-var rename cutover.
Also tidies up pre-existing basedpyright findings in tests/llm/test_model_config.py
(pydantic before-validator dict inputs + descriptor-proxy call).
ruff: clean. basedpyright: 0 errors, 0 warnings. Tests: 153/153 pass across
tests/utils/test_clients.py, tests/utils/test_length_finish_reason.py,
tests/llm/, tests/dialectic/, tests/deriver/.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): finish the src/utils/clients.py → src/llm/ migration
honcho_llm_call_inner now delegates to request_builder.execute_completion
and execute_stream instead of re-implementing backend call scaffolding
inline. The new _effective_config_for_call helper carries per-call kwargs
(temperature, stop_seqs, thinking_budget_tokens, reasoning_effort) onto
the selected ModelConfig — or synthesizes a minimal config for the
test-only callers that pass provider+model directly. max_output_tokens
is zeroed on the effective config to preserve the current
"per-call max_tokens wins" semantic; honoring ModelConfig.max_output_tokens
is a separable correctness concern.
Side effect of routing through the new path: ConfiguredModelSettings'
thinking_budget_tokens validator now fires on synthesized configs.
test_anthropic_thinking_budget was asserting that a sub-1024 budget
propagated to Anthropic — bumped to 1024 to match what Anthropic actually
accepts.
Unified client construction. Promoted the cached client factories in
src/llm/__init__.py (get_anthropic_client, get_openai_client,
get_gemini_client, get_{anthropic,openai,gemini}_override_client) to
public API and added them to __all__. Promoted
credentials._default_transport_api_key → default_transport_api_key.
Deleted the duplicate _build_client and _default_credentials_for_provider
from clients.py; _client_for_model_config now falls through to the
public factories. CLIENTS dict and _get_backend_for_provider stay as the
mockable seam for the ~50 patch.dict(CLIENTS, {...}) test call sites.
Wired operator-configurable Gemini cached-content reuse end-to-end.
PromptCachePolicy moved from src/llm/caching.py into src/config.py so
ModelConfig can reference it as a field without a circular import;
caching.py re-exports the name for existing imports. Added
cache_policy: PromptCachePolicy | None on ConfiguredModelSettings,
FallbackModelSettings, ResolvedFallbackConfig, and ModelConfig.
resolve_model_config, _resolve_fallback_config, and
_select_model_config_for_attempt copy the field through.
honcho_llm_call_inner passes effective_config.cache_policy into
execute_completion / execute_stream, so operators opt in via
e.g. DERIVER_MODEL_CONFIG__CACHE_POLICY__MODE=gemini_cached_content
and the selection actually fires instead of sitting on a dead path.
New regression test test_cache_policy_reaches_gemini_backend asserts the
PromptCachePolicy object reaches the Gemini backend's extra_params.
ruff + basedpyright: clean. Tests: 154/154 pass across
tests/utils/test_clients.py, tests/utils/test_length_finish_reason.py,
tests/llm/, tests/dialectic/, tests/deriver/.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): move all LLM orchestration into src/llm/ and delete clients.py
The 1624-line src/utils/clients.py has been carved up into focused modules
under src/llm/ and deleted. There is now one golden path for LLM
orchestration and no dual entrypoint.
New module layout:
src/llm/
__init__.py thin stable re-export surface
api.py public honcho_llm_call with retry + fallback + tool
loop delegation
executor.py honcho_llm_call_inner (single-call executor); bridges
to request_builder.execute_completion / execute_stream
tool_loop.py execute_tool_loop + stream_final_response, plus
assistant-tool-message and tool-result formatting
runtime.py AttemptPlan dataclass (replaces the loose
ProviderSelection NamedTuple), effective_config_for_call,
plan_attempt, per-retry temperature bump, attempt
ContextVar
registry.py single owner of CLIENTS dict + cached default and
override SDK-client factories + backend/history-adapter
selection + high-level get_backend(config)
conversation.py count_message_tokens, tool-aware message grouping,
truncate_messages_to_fit
types.py HonchoLLMCallResponse, HonchoLLMCallStreamChunk,
StreamingResponseWithMetadata, IterationData,
IterationCallback, ReasoningEffortType, VerbosityType,
ProviderClient
request_builder.py low-level request assembly (ModelConfig → backend
complete/stream); no longer owns credential resolution
credentials.py default_transport_api_key, resolve_credentials
caching.py gemini_cache_store; re-exports PromptCachePolicy
from src.config
backend.py Protocol + normalized result types
history_adapters.py provider-specific assistant/tool message shapes
structured_output.py
backends/ AnthropicBackend, OpenAIBackend, GeminiBackend
handle_streaming_response had no production callers; it is deleted. The
three tests that used it now drive honcho_llm_call_inner(stream=True,
client_override=...) directly, which exercises the same code path the
public API uses.
Dead credential passthrough removed. The ProviderBackend Protocol and
all three concrete backends no longer accept api_key / api_base — those
are baked into the underlying SDK client at registry construction time
and were being del'd everywhere they appeared. request_builder also
stops resolving and forwarding them.
Client construction is unified. The cached default-client factories
(get_anthropic_client, get_openai_client, get_gemini_client) and override
factories (get_*_override_client) are promoted to public API; the
module-level CLIENTS dict populates from them and remains the
patch.dict(CLIENTS, {...}) mocking seam tests rely on. Old duplicate
helpers (_build_client, _default_credentials_for_provider) are gone.
default_transport_api_key is promoted to public.
Application imports now come from src.llm (dreamer, dialectic, deriver,
summarizer, telemetry-adjacent tests). No code imports from
src.utils.clients anywhere in the repo.
ruff: clean. basedpyright: 0 errors, 0 warnings. Tests: 1013/1013 pass
across the entire non-infra test suite (excluding tests/unified,
tests/bench, tests/live_llm, tests/alembic).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(llm): sanitize tool schemas for Gemini's function_declarations validator
Gemini's native-transport function-declarations validator accepts a narrow
subset of JSON-Schema / OpenAPI: type, format, description, nullable, enum,
properties, required, items, minItems, maxItems, minimum, maximum, title.
Anything else — additionalProperties, allOf, if/then/else, $ref, anyOf,
oneOf, $defs, patternProperties — triggers an INVALID_ARGUMENT 400 at call
time.
Our agent tool schemas in src/utils/agent_tools.py use several of those
(additionalProperties: false, allOf + if/then conditionals) because they
were authored for OpenAI strict-mode + Anthropic, which need the richer
vocabulary. GeminiBackend._convert_tools was passing them straight through.
Add _sanitize_schema(): walks the parameters tree and drops unsupported
keywords while preserving semantics for the keywords that hold user data
(properties maps field-name → sub-schema; required / enum are lists of
literals; items is a single sub-schema). Other backends are untouched and
continue to receive the full strict schemas.
Regression tests:
- test_gemini_sanitize_schema_strips_unsupported_keywords: confirms
additionalProperties, allOf + if/then, and $defs are stripped at nested
levels while legitimate fields survive.
- test_gemini_convert_tools_sanitizes_parameters_schema: end-to-end
_convert_tools output has no forbidden keys.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: fix tool calling syntax for gemini
* refactor(llm): normalize defaults, widen OpenAI reasoning-model routing
* chore: fix test
* fix(llm): address post-migration review feedback
* fix(llm): gemini robustness + dreamer specialist ergonomics
* chore: addres review comments
* chore: (docs) unrelease changelog addition
* chore: (docs) merge commit changes
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Erosika <eri@plasticlabs.ai>
* fix: further remove extraneous transactions
* fix: (search) use 2 phase function to reduce un-needed transaction
* fix: refactor agent search to perform external operations before making a transaction
* fix: reduce scope of queue manager transaction
* fix: (bench) add concurrency to test bench
* fix: address review findings for search dedup, webhook idempotency, and bench throttling
* Fix Leakage in non-session-scoped chat call (#526)
* fix: (search) reduce scope for peer based searches
* fix: tests
* fix: (test) address coderabbit comment
* fix: drop db param from deliver_webhook
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: dialectic held connection
* fix: (agent) pre-compute embeddings for agent tools
* fix: (tests) refactor tests to use smaller test db connections
* fix: Embedding client to branch depending on vector store
* fix: reflect dedup-skipped observations in created counts and isolate DB sessions in extract_preferences
* fix: (tests) update tests to match changes
* fix: expunge docs + don't pass in db to query_documents
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: adding strict checking and updating readme
* chore: changelog and version
* fix: Add strict validation to all Python and TypeScript classes
* fix: Address Code Rabit Comments
* fix: duplicate searchQuery param in typescript session.context()
* feat: add created_at, is_active fields, and get_message method
* feat: Add pagintion params to sdk
* fix: Remove lazy initalization behavior from sdks
* fix: Address File Upload Validation, add compatibility shims, address review comments
* chore: Docs updates
* fix: Convert session config from API format in Peer.sessions()
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: Pass all args to Session constructor in Peer.sessions()
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: Preserve createdAt in Peer.refresh(), pass all data in session.peers()
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: changelog for ts
* fix: Review Comments
* fix: Add createAt and to peers call
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
RepresentationManager._query_documents_recent() and
._query_documents_most_derived() do not filter soft-deleted documents,
unlike every other document query function in the codebase. This causes
the deriver's working representation to include documents that are being
garbage-collected.
Refs #444
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>