Add an opt-in `include_evidence` to both chat endpoints. The response then
carries the conclusions and messages the agent read while answering, plus the
tools it called.
Evidence is collated from what the agent accessed rather than reported by the
model. That over-reports -- a conclusion appears because the agent saw it, not
as proof the answer used it -- but it is deterministic, costs no model tokens,
and behaves the same at every reasoning level. Asking the model to cite its
sources fails quietly instead: weaker models produce incomplete or invented
citations, and a sparse citation list is indistinguishable from a sparse
answer. The prefetch already makes the point, feeding explicit conclusions
into the prompt without their IDs, so the model could not cite them if asked.
An accumulator is threaded from the router through the agent into ToolContext,
and read handlers hand it the rows they already loaded. Nothing is re-queried
when the response is built, so evidence inherits the scoping of the reads that
produced it and cannot become a way around a session allowlist.
Three details worth knowing:
- Prefetched conclusions never pass through the tool executor, so they are
captured via a new `documents_out` sink on `search_memory`. On a query that
answers without a tool call they are the whole of what was read.
- Conclusions dedupe by ID. `Representation`'s own deduplication keys on
content and timestamp and ignores IDs, which would collapse distinct
conclusions that happen to read alike.
- Conclusion timestamps are re-stamped UTC. `Representation` strips tzinfo so
observations render compactly into prompts, which would otherwise put naive
timestamps in the API beside timezone-aware message ones.
The two chat routes had byte-identical nested SSE formatters; they now share
one helper, so the terminal event carries evidence on both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: stop top_k=0 from reaching Turbopuffer on message search
HONCHO-19Q: dreamer search_messages passed LLM limit=0 through to
Turbopuffer (top_k must be 1..10000). #970 guarded documents; this
closes the message path and floors tool limits at 1.
* fix: preserve pgvector None sentinel on zero top_k
query_external_vector_document_ids must return None when on the
pgvector path before applying the top_k<=0 empty-list guard.
* fix(deriver): truncate oversize observations so one cannot drop the batch
simple_batch_embed raised ValueError when any input exceeded the per-input
token cap, which failed the entire deriver save when a single observation
was over-length. Add on_oversize="truncate": oversize inputs are embedded
from a token-capped prefix (re-encoded until it fits, with a warning),
preserving one vector per input. Default stays "raise" so existing callers
are unchanged. RepresentationManager opts into truncate.
Also add a live embedding test that fails on main (raise / missing kwarg)
and passes once a mixed short+oversize batch survives.
Refs #569
* fix(deriver): surface failure when all observer saves fail
When every observer's save_representation failed (e.g. embedding retries
exhausted under a sustained 429), the deriver logged the error and returned
normally, so the queue marked the work unit processed with zero documents
saved. Collect per-observer errors and, after telemetry is emitted, raise
RepresentationSaveError when no observer succeeded. Partial failures stay
processed (saved observers must not be discarded) and are recorded via an
additive failed_observer_count on RepresentationCompletedEvent.
Refs #728
* fix(embedding): guarantee truncation progress and truncate on re-embed
The retry slice in _truncate_to_token_limit always recomputed the same
keep count, so a slice whose re-encode grew past the cap could oscillate.
Decrement keep after each unsuccessful retry.
Document re-embed in the reconciler used the default on_oversize="raise",
so one oversize document failed every other document in the batch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: drop ticket ids and shrink comments to one sentence
Comments and docstrings describe current behavior, not the PR that
introduced them. Ticket numbers stay in the commit/PR.
* chore: annotate RepresentationSaveError and assert truncate on re-embed
* fix(embedding): truncate on conclusion create paths and document BPE loop
Storage callers in create_observations (API + agent tools) now pass
on_oversize="truncate" so a single oversize item cannot drop the batch.
Docstring on _truncate_to_token_limit notes why decode/re-encode is load-bearing.
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix: honor DERIVER_DEDUPLICATE in create_observations
The agent-tool path hardcoded deduplicate=True, so DERIVER_DEDUPLICATE=false
could not disable dedup for observations created through this path. Pass
settings.DERIVER.DEDUPLICATE, matching crud/representation.py.
* test: cover deduplicate setting is forwarded in create_observations
* fix: apply session scoping to all working-representation query paths
session_name was only applied to the recent-documents query in
RepresentationManager; the semantic and most-derived paths ignored it,
so limit_to_session leaked cross-session conclusions into perspectives.
- Thread a session allowlist (session_names) uniformly through all
three query paths; pushed down to pgvector and external vector stores
- Accept a list so the upcoming session-allowlist API reuses this path
- Fail closed on an empty allowlist (downstream stores drop empty IN
clauses, which would silently widen scope)
Fixes DEV-1994
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: bare-list membership sugar in the filter DSL
{"session_id": ["s1", "s2"]} is now shorthand for
{"session_id": {"in": [...]}} on regular columns, generically
(peer_id, etc.). JSONB metadata columns are excluded — a bare list
there keeps JSONB containment semantics, unchanged.
Previously a bare list on a regular column compiled to a type-mismatched
equality that matched nothing, so this is strictly additive.
Also translates the same shape in the turbopuffer/lancedb filter
builders, and fixes lancedb dropping empty IN clauses (fail-open) —
an empty membership list now emits an always-false condition.
Groundwork for DEV-1995 (session allowlist via the existing filters
DSL, no new API params)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: session allowlist on dialectic and representation via filters
Adds a constrained 'filters' body to peer.chat and /representation —
the same DSL search and conclusions already accept, supporting only the
session_id key (a session id, a bare list, or {"in": [...]}).
Unsupported keys and shapes are rejected with 422, never silently
ignored. Composes with session_id (must be included in the allowlist
when both are given). Capped at 1,000 sessions per request.
Enforcement is uniform at every recall chokepoint, fail-closed:
- dialectic prefetch + search_memory: conclusion recall restricted to
the allowlist; dream docs (session_name IS NULL) excluded
- message tools (search/grep/date-range/temporal/context/history):
strict intersection of allowlist and observer session membership
- get_reasoning_chain: unavailable under an allowlist (chains traverse
provenance across sessions and cannot be scoped without leaking)
- empty allowlist short-circuits to empty results everywhere
Auth: workspace keys pass the allowlist as-given; peer-scoped JWTs must
be a member of every allowlisted session (403 otherwise), mirroring the
existing single-session check.
Fixes DEV-1995
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: fail closed on empty session allowlist across all filter builders
Empty session allowlists relied solely on the early-return guard in
_get_working_representation_internal. The layers below it were
inconsistent, so a future direct caller (the DEV-1995 allowlist API)
would silently widen scope instead of failing closed:
- _build_filter_conditions used a truthiness check; an empty list was
treated like None and dropped the filter. Now uses `is not None`,
matching the recent/most-derived SQL paths.
- turbopuffer emitted a bare `In []` with undocumented (possibly
fail-open) semantics. Now emits an explicit always-false predicate,
mirroring lancedb's `1 = 0`.
Also extract the duplicated JSONB column tuple in filter.py to a
JSONB_COLUMNS constant.
Tests exercise each fail-closed guarantee at the layer it lives, rather
than masking it behind the early-return guard.
* fix: address tests
* fix(crud): fail closed when session_name is outside the allowlist
search/grep/history helpers scoped to a single session_name ignored the
session_names allowlist entirely — a caller could read a session the
allowlist forbids. The API routes guarded this with a 422, but the
dialectic tools call these CRUD functions directly and bypassed it.
Enforce it at the boundary: return [] when session_name is set and not in
the allowlist, across _semantic_search_messages (covers search_messages +
search_messages_temporal), grep_messages, get_messages_by_date_range,
get_recent_history, and get_observation_context.
Also rename the public param allowed_sessions -> session_names for
consistency with representation.py / chat.py / peers.py; the resolved
intersection keeps its distinct name allowed_session_names.
* fix: test
* fix(scopes): tighten and consolidate session allowlist per review
Addresses review feedback on the session allowlist (DEV-1995).
Behavior changes:
- Auth gate on peers.chat now uses active membership (left_at IS NULL)
via get_peer_session_names(active_only=True), matching the adjacent
is_peer_in_session check on options.session_id. Previously a peer that
had left a session was denied when naming it directly but permitted
when naming it in filters.session_id.
- Scoped conclusion recall is restricted to level == "explicit"
(ALLOWLIST_SAFE_LEVELS). Dream-derived conclusions are stamped with a
single session_name but synthesized across all sessions, so that stamp
can't be scoped on. Applied at all four recall paths. Unscoped recall
is unchanged. Follow-up to give conclusions an authoritative
source-session set is tracked in DEV-2201.
- The allowlist gate checks `is not None` rather than truthiness, so
filters={"session_id": []} reaches it instead of being skipped.
Refactors:
- New crud.message.resolve_session_scope replaces four near-identical
copies of the allowlist-membership intersection. Returns
(allowlist, deny) and never returns an empty list, so the None vs []
distinction that external stores fail open on lives in one tested
place. Takes db=None and opens its own short-lived session only when
distinction that external stores fail open on lives in one tested
place. Takes db=None and opens its own short-lived session only when
an observer lookup is needed, preserving external-lookup-first
ordering on the vector-store path.
- extract_session_allowlist takes must_include, collapsing the
session_id-in-allowlist check duplicated across both peer routes.
- DialecticAgent._select_tools dedupes the two toolset-selection blocks
and drops get_reasoning_chain under an allowlist, rather than paying
for the schema plus a wasted turn to return a refusal.
- Rename session_names -> session_allowlist across crud, agent tools,
dialectic and routes, to remove the one-character ambiguity with
session_name. Internal only; the public filters.session_id surface is
unchanged.
Docs:
- session_allowlist documented across all message and recall entry
points, including the None / [] / populated contract.
- session_name marked deprecated for scoping. Not removed and not
aliased: it also pins the query to one session, bypasses observer
scoping, and drives session-history injection into the dialectic
prompt, so it has no drop-in replacement.
- Note at the Document branch in utils/filter.py that the raw-key
fallback is load-bearing for session scoping.
Tests: 20 -> 39 in tests/test_session_allowlist.py, covering the
peer-scoped JWT gate (member, non-member, left-session, workspace-key
bypass, empty allowlist), the resolve_session_scope tri-state including
the no-DB-checkout path, must_include, and the level narrowing.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix: enforce explicit-document session purity in dedup/merge paths
Audit for DEV-2000 (Scopes RFC prerequisite): explicit-level documents must
stay session-pure so scope memory can be built by copying explicit documents
between collections. Two classes of violation were possible:
- Exact-content and semantic dedup in crud/document.py matched candidates
with no level or session scoping, so an explicit document could be
reinforced by — or soft-deleted in favor of — a same-content document from
a different session or a different level (silently merging cross-session
derivations into one row).
- The generic create_observations tool handler accepted level='explicit'
from agents with no message context (dreamer/dialectic), which would mint
session-less explicit documents.
Enforcement (refuse, never rewrite):
- create_documents refuses explicit documents with a null session_name
- exact dedup keys on (content, level, session-for-explicit); derived levels
keep cross-session consolidation
- is_rejected_duplicate scopes candidate search to the same level, and the
same session for explicit documents
- the create_observations tool rejects explicit-level input outside message
ingestion (deriver) context
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add card_refresh dream type for event-driven peer-card updates
Adds a lightweight dream variant (DEV-2000, Scopes RFC prerequisite) that
runs ONLY the peer-card update — for event-driven refreshes such as scope
membership changes and cold starts:
- DreamType.CARD_REFRESH alongside OMNI; dispatched by process_dream to a
new run_card_refresh_dream orchestration
- CardRefreshSpecialist: restricted to get_recent_observations,
search_memory, and update_peer_card (no observation-mutating tools), with
a low tool-iteration cap of min(6, DREAM.MAX_TOOL_ITERATIONS)
- rebuild=True mode carried in the dream payload: the existing card is NOT
injected into the prompt and the specialist rebuilds it solely from
observations present in the collection (for use after removals)
- enqueue-able via the manual enqueue_dream path (bypasses volume gates);
the work-unit key already embeds the dream type so a card refresh never
collides with a pending omni dream. POST /v3/workspaces/{id}/schedule_dream
accepts dream_type=card_refresh plus the rebuild flag
- card refreshes never advance the omni dream guard pair
(last_dream_at / last_dream_document_count)
- shared PEER CARD prompt section extracted (verbatim) from
DeductionSpecialist for reuse; CallPurpose gains dream.card_refresh
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: fix tests
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(agent_tools): strip display-format "id:" prefix from model-supplied observation IDs
Observations are presented to agents as [id:xxx], and models sometimes
copy the prefix verbatim despite tool-schema instructions to pass the
bare ID. This silently corrupts source_ids provenance on
create_observations_* (broken links stored in document metadata) and
breaks get_reasoning_chain lookups.
Normalize at both entry points. delete_observations is intentionally
not touched here since #746 already covers it.
Only the "id:" prefix is stripped: document IDs are nanoids whose
alphabet includes "-" and "_", so more aggressive cleanup could mangle
legitimate IDs.
Related to #719.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent_tools): strip whitespace remaining after "id:" prefix removal
Addresses CodeRabbit review: defends against "id: xxx" with a space
after the colon, and matches the docstring, which already promised
surrounding-whitespace stripping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: add new cloudevents for api routes
* fix: add total input tokens to RepresentationCompletedEvent
* feat(telemetry): inject honcho_version + emitter health metrics
* feat(telemetry): per-LLM-call event with try/finally emission + sampler
Adds LLMCallCompletedEvent (llm.call.completed) — fires once per provider hit
with full cost-attribution context: transport/provider_label, model, token
counts with cache breakdown, finish_reason, outcome (success or error),
is_final_attempt flag, retry/fallback state, duration, tool-call shape,
streaming flag, and agent correlation (run_id + iteration).
- src/telemetry/events/llm.py: new event class + CallPurpose closed enum
(deriver.representation, dialectic.answer, dream.deduction|induction,
summary.short|long). Resource id includes attempt so multi-attempt retries
in one iteration get distinct deterministic ids.
- src/telemetry/events/base.py: BaseEvent._volume_class ClassVar (default
"ground_truth"); the new event opts into "high_volume".
- src/config.py: TelemetrySettings.HIGH_VOLUME_SAMPLE_RATE (default 1.0).
- src/telemetry/emitter.py: deterministic sampler keyed on run_id (so an
entire agent trace is kept or dropped together). Aggregate envelopes
bypass the sampler. Sampled-out events increment the dedicated counter
separate from buffer_full/send_failed drops.
- src/llm/runtime.py: AttemptPlan gains attempt/retry_attempts/is_fallback
so the executor reads retry state without re-deriving it.
- src/llm/types.py: LLMTelemetryContext dataclass carrying workspace,
call_purpose, run_id, iteration, peer fields. Iteration is mutable so
the tool loop can set it per inner call.
- src/llm/executor.py: honcho_llm_call_inner wraps the backend call in
try/finally — emits on success AND on exception, with is_final_attempt
computed from AttemptPlan. Stream path emits a was_stream=True placeholder
(token totals deferred until streaming completion is wired through).
Telemetry failures swallowed.
- src/llm/api.py: threads telemetry kwarg through all 4 signatures into
both honcho_llm_call_inner and execute_tool_loop.
- src/llm/tool_loop.py: _telemetry_for_iteration helper copies the caller
context with iteration set per call — covers both the normal iteration
loop AND the max-iteration synthesis call (iteration N+1).
Tests cover success/error emission, sampler trace-coherence (same run_id →
same decision), volume_class enforcement, unknown call_purpose tolerance,
provider_label inference, and telemetry failure isolation. 378/378 pass.
* feat(telemetry): emit agent.iteration on every LLM response + synthesis
AgentIterationEvent was defined but never emitted on this branch. Phase 2
wires it up in execute_tool_loop so every LLM call inside an agentic loop
produces one event — including the no-tool terminating iteration and the
max-iteration synthesis call — and threads LLMTelemetryContext from dialectic
and dreamer specialists down through honcho_llm_call.
- src/telemetry/events/agent.py: AgentIterationEvent opts into
_volume_class="high_volume" so the Phase 1 sampler throttles it.
- src/llm/tool_loop.py: _emit_agent_iteration() helper fires once per
honcho_llm_call_inner response, BEFORE the no-tool early return so the
terminating iteration is counted. A second emission fires for the
max-iteration synthesis call BEFORE final_response is mutated with
cumulative totals (otherwise the per-iteration counts would double-count).
Emission is defensively skipped when telemetry context lacks run_id /
agent_type / parent_category / workspace_name; emit failures are swallowed.
- src/dreamer/specialists.py: BaseSpecialist.run passes LLMTelemetryContext
with parent_category="dream", agent_type=self.name, observer/observed,
call_purpose=f"dream.{self.name}".
- src/dialectic/core.py: _telemetry_context() builds a shared context for
both answer() and answer_stream(), using self._run_id (always set) +
workspace + observed peer.
Tests cover fresh-copy semantics, per-iteration vs terminating emission,
defensive skip cases, telemetry-failure isolation, and volume_class. 408/408
pass across telemetry + llm + utils + dreamer + dialectic.
* feat(telemetry): agent.tool.call.completed event + ToolResult metadata
Adds the missing generic per-tool-call event so read-only tools (search_*,
get_recent_history, get_observation_context, etc.) and the four existing
state-change tools all produce a telemetry record. Built on a new internal
ToolResult(content, metadata) contract so handlers can surface
search-specific fields (top_k/used_embedding/query_tokens/results_count)
to Phase 3 and create/delete counts to Phase 5's specialist rollups.
- src/telemetry/events/agent.py: AgentToolCallCompletedEvent at v1 with
_volume_class="high_volume". Resource id = {run_id}:{iteration}:{tool_call_seq}
so two calls to the same tool in one iteration don't collide
deterministic ids and get dedup-dropped downstream.
- src/utils/types.py: ToolResult dataclass; two new ContextVars
(_current_tool_call_seq + _last_tool_metadata) so tool_loop and the
execute_tool closure can communicate per-call telemetry without changing
the public Callable[[str, dict], Any] signature.
- src/utils/agent_tools.py: execute_tool times handlers, unwraps ToolResult,
publishes metadata, emits the event. Handlers updated to ToolResult
where useful: create/delete observations, update_peer_card, search_memory,
search_messages. Other handlers continue to return str.
- src/llm/tool_loop.py: set_current_tool_call_seq before each executor call;
read get_last_tool_metadata after and stash on all_tool_calls[i] for
Phase 5 rollups.
Tests cover ToolResult str-likeness, ContextVar round-trip, full-context
emission with search metadata, resource-id disambiguation, defensive skip
cases, telemetry isolation, truncation metadata, volume_class. 420/420 pass.
* feat(telemetry): RepresentationCompletedEvent v2 token breakdown + tool-less truncation
Bulks out the deriver's per-batch telemetry without bumping the event schema
version. New additive fields capture the full token breakdown (queued vs.
extra-context vs. scaffold), the cap configuration (batch_max_tokens,
max_input_tokens, was_flush_enabled), real cap-hit flags, and observer
fanout. `input_tokens` stays unchanged as the queued-message-tokens billing
key Xatu's Stripe meter reads.
The big enabler: src/llm/api.py now actually enforces max_input_tokens on
the tool-less LLM path. Before this, the deriver passed the kwarg but the
path silently dropped it — so the configured cap was advisory and
hit_input_token_cap couldn't be measured. Phase 4 wires truncation through
the same truncate_messages_to_fit helper the tool loop uses and surfaces
input_was_truncated on HonchoLLMCallResponse.
- src/telemetry/events/representation.py: 12 additive fields, schema_version
stays at 2.
- src/llm/types.py: input_was_truncated on HonchoLLMCallResponse.
- src/llm/api.py: tool-less path truncates messages before dispatch, flips
input_was_truncated on the response when clamping occurs. Split into
Literal[True]/Literal[False] branches for typecheck.
- src/deriver/queue_manager.py: QueueBatchResult dataclass replaces the
3-tuple return from get_queue_item_batch; carries hit_batch_token_cap
(computed from cumulative token sum vs cap), was_flush_enabled snapshot,
and batch_max_tokens. Worker loop unpacks + forwards.
- src/deriver/consumer.py: process_representation_batch gains the three
flag kwargs and forwards.
- src/deriver/deriver.py: derives the breakdown fields locally, populates
the new fields on emit, sources hit_input_token_cap from
response.input_was_truncated.
Tests cover schema stability, defaultable fields, input_tokens semantic
preservation, cap-hit flag round-trip, model_dump completeness, and
HonchoLLMCallResponse.input_was_truncated mutability. Existing
test_queue_processing.py tests updated for QueueBatchResult and mock
process_representation_batch signature. 479/479 pass.
* feat(telemetry): DreamRunEvent v2 scheduler reasons + DreamSpecialistEvent v2 rollups
Bumps both dream events to v2 with additive fields. DreamRunEvent gains
scheduler context (threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold / dream_type / enabled_types_count) threaded through the
dream queue payload — the two scheduler gates stay as separate fields rather
than collapsing into one trigger_reason, preserving the WHY-vs-WHEN
semantics. DreamSpecialistEvent gains denormalized rollups
(created_observation_count / deleted_observation_count / peer_card_updated /
search_tool_calls_count) sourced from Phase 3's ToolResult.metadata so the
counts reflect observation truth, not call truth.
- src/telemetry/events/dream.py: schema_version → 2 for both events; new
fields all defaultable so older producers still construct valid events.
- src/utils/queue_payload.py: DreamPayload + create_dream_payload accept
threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold.
- src/dreamer/dream_scheduler.py: check_and_schedule_dream computes the two
reasons at decision time and threads them through schedule_dream →
_delayed_dream → execute_dream → enqueue_dream.
- src/deriver/enqueue.py: create_dream_record / enqueue_dream gain the
kwargs and persist on the queue payload.
- src/dreamer/orchestrator.py: process_dream unpacks the payload; run_dream
accepts the kwargs and stamps them on DreamRunEvent.
- src/dreamer/specialists.py: BaseSpecialist.run walks response.tool_calls_made
and sums ToolResult.metadata.created_count / .deleted_count, sets
peer_card_updated, counts search-tool calls by name.
Tests cover schema_version bumps, defaultable Phase 5 fields,
threshold-vs-delay semantics, observation-vs-call-count rollup distinction,
and DreamPayload round-trip. Existing tests updated for the schema bump
and the new enqueue_dream kwargs. 488/488 pass.
* feat(telemetry): AgentToolSummaryCreatedEvent v2 token breakdown
Bumps schema_version to 2 and adds three additive breakdown fields so
analytics can answer "how much of a summary call's cost was the previous-
summary rollup vs. the new messages vs. the scaffold instructions".
- src/telemetry/events/agent.py: previous_summary_tokens, message_tokens,
prompt_scaffold_tokens added with sensible 0 defaults. input_tokens
retains its current semantic (provider-side LLM tokens) — the plan's
proposed `provider_input_tokens` was omitted because input_tokens
already serves that purpose and a duplicate would fork queries.
- src/utils/summarizer.py: emit now populates the three new fields from
values already in scope (messages_tokens, previous_summary_tokens,
prompt_tokens). Hoisted prompt_tokens calculation out of the
is_fallback conditional so both the save-summary path and the emit
share one binding — basedpyright couldn't prove the sibling-scope
binding was safe, and the compute is cheap + idempotent.
Tests cover schema bump, defaultable fields, input_tokens semantic
preservation, first-summary edge case, and breakdown round-trip.
493/493 pass.
* feat(telemetry): embedding.call.completed event + call-purpose ContextVar
Adds the final piece of cost-attribution telemetry: per-embedding-call
events covering every provider hit (single + batch + retry attempts).
Embedding calls are real provider spend that was invisible before this
phase; search-heavy paths (dialectic agentic) can produce more embedding
calls than LLM calls, so the new event participates in the shared
HIGH_VOLUME_SAMPLE_RATE.
- src/telemetry/events/llm.py: EmbeddingCallCompletedEvent at v1 with
_volume_class="high_volume". EmbeddingCallPurpose closed enum
(search_memory / search_messages / create_observations / vector_sync /
summary / message_create). Resource id = run:purpose:provider:model:input_count
so per-iteration calls in one agentic run don't collide.
- src/utils/types.py: _embedding_call_purpose ContextVar plus
@contextmanager wrapper. Nesting-safe via ContextVar.reset(token).
Callers wrap embedding-driving operations in
`with embedding_call_purpose("search_memory"): ...` — no changes to
the embedding client signature.
- src/embedding_client.py: _emit_embedding_call wraps each provider hit
with try/finally so success AND error paths emit. Errors propagate
unchanged. Each retry attempt of _process_batch emits its own event.
Unknown call_purpose slugs drop to None (validation against the enum
happens at emit time, not at context-manager-set time).
- src/utils/agent_tools.py: search_memory / search_messages /
search_messages_temporal / create_observations (batch + fallback) all
tag their embedding calls.
- src/crud/representation.py: save_representation tags with
CREATE_OBSERVATIONS; get_working_representation precompute tags with
SEARCH_MEMORY.
- src/crud/message.py: create_messages batch embed tags with
MESSAGE_CREATE; search_messages/temporal fallback tags with
SEARCH_MESSAGES.
Tests cover event shape, enum closure, ContextVar nesting/exception
cleanup, wrapper success+error emission, unknown-purpose graceful
fallback, telemetry-failure isolation. 550/550 pass across the full
telemetry+llm+utils+dreamer+dialectic+deriver+crud test set.
* chore: fix tests
* fix(telemetry): address review findings on stream events, context propagation, and cap detection
Five findings from a post-Phase-7 review (one resolved by the merge from
main, four addressed here):
- src/llm/executor.py: stream-path LLMCallCompletedEvent now fires AFTER
the stream is set up and drained (or on exception), with real duration
and accurate outcome. Previously the event was emitted before
execute_stream() ran and was always recorded as outcome="success" with
duration_ms=0, which silently masked stream-setup and stream-drain
failures. Wrapping the async generator in try/finally surfaces the real
outcome; token counts stay 0 because we still don't have them at stream
end (aggregate envelopes carry totals).
- src/deriver/deriver.py + src/utils/summarizer.py: deriver and summarizer
LLM calls now thread LLMTelemetryContext into honcho_llm_call. Before
this, the closed CallPurpose enum had DERIVER_REPRESENTATION /
SUMMARY_SHORT / SUMMARY_LONG slugs but those production call sites
didn't actually pass `telemetry=`, so their LLMCallCompletedEvents lost
workspace_name, parent_category, and call_purpose. summarizer threads
workspace_name through _create_and_save_summary → _create_summary →
create_short_summary / create_long_summary.
- src/utils/types.py + src/embedding_client.py: embedding_call_purpose
ctx manager now accepts workspace_name and run_id kwargs, backed by
two new ContextVars. EmbeddingCallCompletedEvent's publisher reads
both via get_embedding_workspace_name / get_embedding_run_id so
embedding events carry workspace and run correlation. All call sites
updated: search_memory / search_messages / search_messages_temporal /
_handle_create_observations_impl pass ctx.workspace_name +
ctx.run_id; create_observations standalone and create_messages pass
workspace_name; RepresentationManager.save_representation and
get_working_representation pass self.workspace_name.
- src/deriver/queue_manager.py: hit_batch_token_cap detection rewritten.
Previously summed kept-rows' token_count and checked against
batch_max_tokens, but the SQL filter `cumulative_token_count <= cap`
guarantees kept rows stay under the cap, so the flag almost never
fired. Now uses two follow-up queries: total token_count across the
included id range + EXISTS check for any session message past the
last-kept id. Both true → cap was actually binding.
(The fifth finding — deriver scaffold-token computation needing
estimate_deriver_prompt_tokens(custom_instructions) — was resolved by
the merge from main; the Phase 4 emit at src/deriver/deriver.py:283
already sources prompt_scaffold_tokens from the wrapped helper.)
567/567 telemetry+llm+utils+dreamer+dialectic+deriver+crud tests pass.
ruff + basedpyright clean.
* chore: ruff linting
* chore: clean AI generated comments references specs
* fix: address coderabbit changes
* fix(telemetry): address remaining PR review findings
Six findings from the PR 637 telemetry review batched into one commit.
- src/llm/executor.py + src/embedding_client.py: asyncio.CancelledError
now surfaces as outcome="cancelled" on both stream and sync paths,
distinct from "error". Client disconnects mid-stream and server
shutdowns are normal control flow and should not feed error-rate
alerting. LLMCallCompletedEvent and EmbeddingCallCompletedEvent
outcome Literal extended; docstrings + tests cover the new state.
- src/utils/types.py + src/llm/tool_loop.py: new iteration_scope()
context manager captures and resets the four per-tool-loop
ContextVars (_current_iteration, _current_tool_call_seq,
_current_provider_tool_call_id, _last_tool_metadata). Applied as a
typed decorator to execute_tool_loop so back-to-back loops in the
same asyncio Task (worker batches, tests) don't observe stale state.
- src/telemetry/events/api.py + src/routers/messages.py:
MessageCreatedEvent schema v1 → v2. Added required last_message_id
(nanoid public_id of the trailing message); get_resource_id now keys
on it instead of message_count, eliminating the collision case where
two same-size batches in the same session+source produced identical
event ids. message_count stays on the body for analytics.
- src/deriver/queue_manager.py: hit_batch_token_cap now computed from
the FINAL post-config-filter batch. Previously the flag used the
pre-filter messages_context[-1].id, which produced false positives
when _resolve_batch_configuration trimmed the trailing queue item —
telemetry reported a cap-hit when the actual returned batch was
short for unrelated reasons. Cap-detection block moved inside the
async with after the filter; no extra DB connection.
- src/config.py + src/telemetry/emitter.py: documented the
HIGH_VOLUME_SAMPLE_RATE orphan trade-off (rate<1.0 keeps aggregates
but drops children, so JOIN ON run_id queries see partial traces).
Behavior unchanged — rate defaults to 1.0.
- src/deriver/deriver.py: WARNING-level invariant logs when
response.input_tokens < messages_tokens (provider tokenization
drift) or prompt_scaffold_tokens <= 0 (estimator silent failure).
Best-effort — telemetry never bleeds into the deriver path
* fix(telemetry): stream retry, embed attempts, truncation, dedup
Address remaining audit findings on the cloudevents PR:
- Stream setup now runs inside the awaited honcho_llm_call_inner so
tenacity's retry wrapper in stream_final_response catches transient
setup failures (rate-limit, auth, network). Previously the returned
generator deferred execute_stream until first iteration — outside
the retry wrapper — crashing the request and bypassing telemetry.
- Embedding _emit_embedding_call gains an is_final_attempt parameter;
_process_batch threads the real retry index so dashboards stop
conflating one-shot, mid-retry, and exhausted-retry calls.
- _truncate_tool_output returns (text, original_chars, was_truncated)
and a new _maybe_truncated_result helper wraps in ToolResult when
truncation happens. Five handlers migrated. AgentToolCallCompletedEvent
fields was_truncated and result_chars_before_truncation are now
populated instead of always None/False.
- execute_tool_loop tracks any_iteration_truncated and stamps
input_was_truncated on the final response (both HonchoLLMCallResponse
and StreamingResponseWithMetadata). Dialectic now reports
hit_input_token_cap correctly.
- GetContextEvent.get_resource_id uses empty-string sentinel instead
of literal "none" so a peer named "none" can't collide with absent.
- generate_event_id folds honcho_version into the deterministic id so
same logical event from different deploys produces distinct ids.
* fix(telemetry): address audit findings across LLM/embed/event paths
Three rounds of telemetry audit findings, grouped by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). Previously the
inner generator deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation populated by _truncate_tool_output via
a new _maybe_truncated_result wrapper; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose context manager accepts parent_category.
- 4 new EmbeddingCallPurpose enum values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Queue batch cap detection (P2.1)
- hit_batch_token_cap keys on the pre-config-filter SQL boundary so the
"kept=900 of 1000 cap, next=300 excluded by cap" case reports True
while still avoiding the config-filter false positive.
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3) — top_k,
used_embedding, embedding_query_count, query_tokens, results_count.
Tests: stream-setup retry, stream-retry attempt sequence, post-stream
output_tokens write-back, is_final_attempt matrix, truncation E2E,
tool-loop hit_input_token_cap propagation, honcho_version in event id,
GetContextEvent disambiguation, queue_items_cleaned round-trip.
* fix(telemetry): address audit findings across LLM/embed/event paths
Four rounds of telemetry audit findings (initial + 3 follow-ups), grouped
by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). The inner
generator previously deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Queue batch cap detection
- hit_batch_token_cap previously required total_in_range >= cap, which
produced false negatives whenever the kept range didn't fully exhaust
the budget. Replaced with a pre-config-filter SQL boundary check
(P2.1), then further refined to a queue-item boundary comparison
(Fix 14) so trailing-context trimming doesn't false-negative either.
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation now populated by _truncate_tool_output
via _maybe_truncated_result; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11). except BaseException
catches cancellations too (Fix 16).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12), and uses ValidationException for unsupported types per
project guideline (Fix 17).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose accepts parent_category.
- 4 new EmbeddingCallPurpose values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Reconciler no longer holds DB session during embedding (Fix 15)
- _sync_documents and _sync_message_embeddings refactored into
three phases per CLAUDE.md guideline: fetch+detach in a small DB
scope, external embedding call without DB locks, writes in a fresh
short-lived DB scope. New _apply_*_sync helpers; orchestrators
expunge ORM objects before invoking. Vector store upsert + sync_state
updates stay in the apply phase together.
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3).
- Dialectic.prefetched_conclusion_count uses Representation.len() so
inductive + contradiction observations count too (Fix 10).
* fix(telemetry): orchestrator emit + review feedback
Three more rounds of audit findings + inline PR review, grouped:
Orchestration / emit reliability
- run_dream wrapped in try/finally so DreamRunEvent always emits, even
on unexpected exceptions including CancelledError (`finally` still
runs while cancellation propagates). Specialist except clauses
broadened from SpecialistExecutionError (never raised in src/) to
Exception so provider/DB/tool failures are recorded with
deduction_success=False / induction_success=False instead of crashing
past the emit.
- BaseSpecialist.run() telemetry state initialization + try/finally
hoisted above the preflight phase (peer lookup, peer-card preload,
create_tool_executor, get_model_config, prompt construction) so
preflight failures emit DreamSpecialistEvent(success=False) instead
of being dropped on the floor.
- Reverted the Round-4 _sync_documents / _sync_message_embeddings
phase split. The split introduced a race: rows were released from
FOR UPDATE SKIP LOCKED before the embed call, allowing two workers
to claim and clobber the same batch. Long-held DB transaction
restored (pre-existing CLAUDE.md violation accepted as a deliberate
trade-off; proper fix requires a claim/in_flight migration tracked
separately).
Schema + naming (PR-internal — none of these have shipped)
- threshold_reason → trigger_reason on DreamRunEvent, DreamPayload, and
every emit/scheduler/router/test call site (~45 src + 21 test lines).
Name now accurately reflects the field's role across "manual",
"surprisal", and "document_threshold" values.
- MessageCreatedEvent reset to schema v1 (was internally bumped to v2
for last_message_id but never shipped at v1 — downstream sees it
for the first time at merge).
- DreamSpecialistEvent gains created_counts_by_level /
deleted_counts_by_level: dict[str, int] keyed on the closed
level taxonomy. Per-tool-call events use list[str] (≤10 items),
but specialist runs aggregate 20+ — dict keeps emissions compact.
- QueueBatchResult marked frozen=True.
Per-call embedding attribution
- Agent tool embedding_call_purpose wraps for search_memory,
search_messages, search_messages_temporal, create_observations now
driven embedding cost rolls up under the right workflow.
- create_observations() signature gains parent_category kwarg
(mirrors existing run_id pattern).
Manual dream scheduling
- Manual /schedule_dream route now passes trigger_reason="manual" and
delay_reason="immediate". Previously both arrived as null in
DreamRunEvent, breaking analytics joins.
Queue-batch SQL perf
- next_exists_check folded into the main CTE query via
bool_or(cumulative_token_count > batch_max_tokens) OVER () in a
nested subquery. Cap detection is now one roundtrip per batch
instead of two.
Code/doc cleanup
- representation.py docstring uses generic "downstream metering key"
language (was "Xatu's Stripe meter"). bench runner --base-url help
uses a generic example host (was "groudon.fly.dev"). Public-facing
code/docs shouldn't reference internal service names.
Tests added for: orchestrator failure-path DreamRunEvent emission,
specialists preflight try/finally coverage, manual-dream
trigger_reason/delay_reason round-trip, dict-rollup accumulation across
multiple tool calls in a specialist run, CTE-fold one-roundtrip
behavior. Full Python suite passes (1236).
* fix(telemetry): correctness + attribution + emitter robustness
- Dreamer iteration count: read response.iterations directly so
one-shot runs no longer report iterations=0 and tool-using runs
include the terminal/synthesis LLM call.
- RepresentationCompletedEvent.observer_count counts successful
saves, not attempts.
- search_memory empty-memory fallback reports the snippet count when
message context is returned (was always 0).
- Wire parent_category through every embedding emit path: message
create (api), save_representation (representation), per-observation
fallback (caller-supplied), and the peer/session context routes
(api). get_working_representation accepts parent_category and
embedding_purpose so the internal fallback embed lands in the same
analytics bucket as the route-level precompute even when the
precompute is suppressed.
- BatchItem carries token_count so _process_batch reuses chunk-prep
counts instead of re-encoding every chunk for the telemetry proxy.
- Drop vestigial EmbeddingCallCompletedEvent.batch_size (always ==
input_count).
- Emitter: release the lock during HTTP send so a failing endpoint's
retry+backoff (~36s worst case) doesn't block other flushers;
edge-trigger the 80%-capacity warning so sustained backpressure
doesn't flood logs; defer event_id generation past the high-volume
sampler for events with run_id so sampled-out children don't pay
the sha256; harden emit() against sync callers with no running
loop; track threshold-flush tasks so shutdown() drains in-flight
sends before closing the HTTP client.
* fix(telemetry): tool cancellation emit, nanoid run_ids, version unification
- execute_tool: wrap post-work in finally so AgentToolCallCompletedEvent
fires on CancelledError; explicit handler sets is_error/result_str
before re-raising.
- run_id: replace str(uuid.uuid4())[:8] with generate_nanoid() across
dialectic/dreamer/specialists; matches project-wide nanoid convention.
- Bump _schema_version on events touched by run_id widening:
DialecticCompletedEvent v1→v2 (also covers hit_input_token_cap field),
AgentIterationEvent v1→v2, AgentToolConclusionsCreatedEvent v1→v2,
AgentToolConclusionsDeletedEvent v2→v3, AgentToolPeerCardUpdatedEvent
v1→v2.
- Unify honcho_version: single HONCHO_VERSION constant in src/_version.py
read from pyproject.toml (importlib.metadata fallback). Drop
TELEMETRY.HONCHO_VERSION setting. Use the constant for the FastAPI app
version (no more hardcoded "3.0.6") and for emitter body injection.
- Delete 17 tautological per-event test_schema_version methods; the
parametrized contract test still enforces version >= 1 across all events.
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix(deriver): ignore blank observations before embedding
* Address PR review on observation normalization
* Harden mock await arg access in tests
* Unify blank observation filtering across tool paths
* Move soft-delete query test back to fixture class
* fix: further remove extraneous transactions
* fix: (search) use 2 phase function to reduce un-needed transaction
* fix: refactor agent search to perform external operations before making a transaction
* fix: reduce scope of queue manager transaction
* fix: (bench) add concurrency to test bench
* fix: address review findings for search dedup, webhook idempotency, and bench throttling
* Fix Leakage in non-session-scoped chat call (#526)
* fix: (search) reduce scope for peer based searches
* fix: tests
* fix: (test) address coderabbit comment
* fix: drop db param from deliver_webhook
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: dialectic held connection
* fix: (agent) pre-compute embeddings for agent tools
* fix: (tests) refactor tests to use smaller test db connections
* fix: Embedding client to branch depending on vector store
* fix: reflect dedup-skipped observations in created counts and isolate DB sessions in extract_preferences
* fix: (tests) update tests to match changes
* fix: expunge docs + don't pass in db to query_documents
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: Add bounds to gemini client
* fix: Prevent empty summaries from being saved to DB (HONCHO-M7)
Raise LLMError on blocked Gemini responses (SAFETY, RECITATION, etc.)
so retry/backup-provider logic triggers. Treat empty LLM responses in
the summarizer as fallback instead of persisting empty strings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: Summary Eval via Locomo
* fix: Code Rabbit Comments
* fix: Code Rabbit Comments
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>