* fix(deriver): eliminate create_documents deadlock and stop silently burning batches on transient errors
Two concurrent work units writing the same (workspace, observer, observed)
collection deadlocked on times_derived reinforcement UPDATEs issued in
batch order (DEV-1975, 682 events in 90 days). The deadlock was swallowed
per-document, the loop cascaded PendingRollbackErrors against the dead
session, the whole batch was lost, and the queue item was marked processed.
- serialize writers per collection with a transaction-scoped advisory lock
(pg_advisory_xact_lock + SET LOCAL lock_timeout), skipped for insert-only
batches; covers all three row-lock sites in one move
- hoist external-vector-store dup-candidate resolution ahead of the first
DB statement so the lock's critical section contains no network calls
- abort the batch on SQLAlchemyError instead of continuing through an
aborted transaction; per-document skip semantics kept for non-DB errors
- classify transient errors (new src/utils/retryable_errors.py) and retry
them via a bounded in-process counter instead of marking items errored
* fix(deriver): replace create_documents advisory lock with id-ordered row locks
Advisory locks are database-scoped and would serialize every writer to a
collection, including across Groudon tenants that share names. Collect
reinforcement and replace ops during the loop, lock target rows with
SELECT ... ORDER BY id FOR UPDATE, then apply. populate_existing reloads
times_derived so a prefetched identity-map row cannot lose a concurrent
increment.
* fix(deriver): harden create_documents candidate hoist and test isolation
Skip empty embeddings on the external-store path, isolate per-document
resolve failures, and keep replacement times_derived in the in-batch
ledger. Patch get_external_vector_store in the hoist test and cover
in-loop SQLAlchemyError abort.
* fix(deriver): address CodeRabbit findings on create_documents deadlock fix
- Distinguish external resolve failure ([] skip) from pgvector fallback (None)
so _semantic_dup_decision never re-enters external I/O under an open session
- Bound external candidate hoist concurrency with a semaphore
- Map in-loop IntegrityError to ValidationException for a uniform contract
- Persist transient retry attempts on the oldest unprocessed queue item so
every deriver instance shares one MAX_RETRYABLE_ATTEMPTS budget
- Cover resolve-failure skip and multi-manager reclaim of the retry budget
* fix(deriver): harden retry metadata cleanup and stale reinforce fallback
- Strip _retry_attempts from payloads in the same transaction as
mark_queue_items_as_processed / mark_queue_item_as_errored
- Clear shared retry metadata only after a successful terminal mark
- On reinforce, if the locked target is gone or soft-deleted, insert the
incoming document instead of dropping it
- Skip pgvector semantic lookup when embedding is empty so query_documents
cannot embed under an open session
* fix(deriver): address review on deadlock retry and row-lock apply
Strip _retry_attempts before payload validation so non-representation
tasks are not burned as extra_forbidden. Re-raise retryable observer
save errors after telemetry so the queue actually retries. Skip
same-batch reinforce fallbacks after a replace. Revert unordered
FOR UPDATE on mark processed/errored and drop post-commit retry
cleanup from the success path.
* fix: add test and simplify queue query
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix(deriver): strip NUL bytes from model-generated observations
Postgres rejects NUL (0x00) in text columns and in jsonb strings. API
ingress has always stripped it from user-supplied content, but the
deriver's own output did not go through any equivalent: a model can emit
a \u0000 escape in its tool-call arguments, which the JSON parser decodes
into a real NUL byte. Seen in production when models transcribe shell
output (`tr '\x00' '\n'`) or Windows paths (`c:\<NUL>users\amal`).
The NUL reached the exact-content dedup pre-fetch in create_documents as
a bind parameter, so the query raised DataError before any row was
written and the whole batch for that observer was dropped.
Strip in _normalized_observation and _normalized_observation_input --
the points that already normalize text for persistence and embedding --
so the embedded text matches the stored text. premises and sources are
covered too, since they ride along in internal_metadata. The emptiness
check now runs after normalization, because str.strip() does not remove
NUL and all-NUL content would otherwise be stored as an empty string.
DocumentCreate.content gets a mode="before" validator as a backstop for
callers that bypass those paths; running before the length constraint
makes all-NUL content fail min_length rather than silently empty out.
The NUL helpers move out of schemas/api.py into utils/sanitization.py as
a single recursive strip_nul, so ingress and internal paths share one
implementation. It is overloaded to keep str -> str for the callers that
chain .strip(), and passes None through so optional fields need no guard.
Fixes HONCHO-4XZ
* fix: broaden nul strip check
* chore: code simplification
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix: stop top_k=0 from reaching Turbopuffer on message search
HONCHO-19Q: dreamer search_messages passed LLM limit=0 through to
Turbopuffer (top_k must be 1..10000). #970 guarded documents; this
closes the message path and floors tool limits at 1.
* fix: preserve pgvector None sentinel on zero top_k
query_external_vector_document_ids must return None when on the
pgvector path before applying the top_k<=0 empty-list guard.
`get_observation_context` resolved scope by fetching every session name the
observer has a membership record in, then expanding that list into
`session_name IN (...)` twice in one statement — once in the CTE and once in
the outer select. That puts psycopg's 65535-bind-parameter ceiling at roughly
32,765 sessions, and the count only ever grows: the loose membership
definition (`active_only=False`) counts sessions the peer has since left, so
leaving a session does not shrink the scope. A workspace with tens of
thousands of sessions for one peer produced a statement the driver could not
serialize at all.
Two new helpers in `crud.message` express the observer half as a correlated
EXISTS over `session_peers`. Scope now costs two bind parameters regardless of
membership size, and the membership query disappears (two round trips become
one). The `session_peers` primary key is `(workspace_name, session_name,
peer_name)`, so the correlated probe is an exact-match index hit.
The caller-supplied allowlist stays an IN clause — it is route-capped at 1000
entries and carries none of the unbounded-growth risk. `resolve_session_scope`
is left in place: three other callers still need the materialized list,
including `_search_messages_external`, which sends session names to the vector
store as a filter payload and cannot take SQL.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Callers wrapped every ValueError from the embedding client in a
"exceeds maximum token limit" message, so provider and configuration
failures (dimension mismatch, empty response, upstream error) surfaced
to users as though their input were too long.
Add EmbeddingTokenLimitError, raised only by the pre-flight token checks
in embed() and simple_batch_embed(), and narrow the remaps in search.py,
agent_tools.py, document.py and representation.py to catch it. It
subclasses ValueError so existing broad handlers keep working.
Both simple_batch_embed() remap sites pass on_oversize="truncate" and so
could never raise a token-limit error at all; their handlers only ever
mislabelled provider failures.
Fixes#568
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add workspace-level chat (DEV-1326)
POST /v3/workspaces/{workspace_id}/chat: agentic dialectic over the whole
workspace instead of a single (observer, observed) pair. Salvaged from
plastic-labs/honcho#373 and re-grown on today's DialecticAgent:
- WorkspaceDialecticAgent subclasses DialecticAgent via four new seams
(_get_tools, _create_tool_executor, _prefetch_intro, _trace_name) instead
of a base-class extraction; observer/observed use empty-string sentinels.
- Routing-accelerated prefetch: workspace stats + top-5 active peers with
their self peer-cards (pure DB, ~7ms measured) so routing-obvious queries
resolve without a discovery tool round.
- Observation search stays pair-scoped (matches per-pair vector namespaces;
avoids workspace-flat top-k dilution): search_memory/get_peer_card take
observer/observed as tool arguments, with pair attribution in results.
- workspace_chat / workspace_chat_stream orchestrators, WorkspaceChatOptions
schema (scope param seam left for the #897 scopes facade), SSE streaming,
structured output via response_format.
- crud: get_workspace_stats, get_active_peers; format_documents_with_attribution.
- SDKs: Python Honcho.chat/chat_stream + HonchoAio mirrors; TypeScript
honcho.chat/chatStream.
- 46 tests (route, orchestrator preflight, tool handlers, executor routing,
attribution formatting) + unified test cases + docs.
Co-Authored-By: doria <93405247+dr-frmr@users.noreply.github.com>
Co-Authored-By: Benjamin McCormick <docterformer@protonmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: type SSE stream wrapper as AsyncIterator (basedpyright)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: silence unused db_session fixture warnings (basedpyright failOnWarnings)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: drop docs changes from this PR (defer to follow-up)
Restores docs/v3/documentation/features/chat.mdx to main's version. This
also puts back the peer-chat Structured Outputs section (#896) that the
workspace-chat commit removed as a rebase artifact.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: workspace message tools deny-all under rebased session scoping
The #882 rebase changed the unscoped-observer contract from falsy to
'observer is None': resolve_session_scope looked up the workspace
executor's observer='' sentinel as a real peer with no session
memberships and denied every workspace-flat message read (search, grep,
date-range, temporal, observation context) whenever no session was
pinned — the primary workspace-chat shape. Normalize the sentinel to
None at the five read-handler crud boundaries and add regression tests
that run the tools unpinned (verified to fail without the fix).
Also from review:
- wrap the workspace prefetch in the same degrade-to-None protection
the base agent has (an overview query error no longer 500s the
request or kills the SSE stream after headers)
- thread session_allowlist through create_workspace_tool_executor so
the agent-level allowlist seam is honored end to end when scopes
(#897) wire it up; allowlisted grep is covered by a test
- deterministic name tie-break in get_active_peers ordering
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: SDK response_format parity, shared query sanitizer, annotations
- TS SDK: WorkspaceChatParams gains response_format; _workspaceChat/
_workspaceChatStream consume the shared interface instead of inline
duplicates; chat/chatStream expose responseFormat.
- Consolidate the three identical sanitize_query validators into one
NulStripped annotation.
- workspace_chat_stream: return annotation + full docstring.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: fold active peers into workspace stats; trace + query bounds
- Merge get_active_peers into get_workspace_stats (one discovery round
instead of two); minimal loadout keeps a discovery tool via the merged
stats tool. Fixed top-10 by recent activity; deeper discovery routes
through search_messages.
- get_active_peers CRUD now aggregates over a trailing 90-day window so
the chat-path prefetch never scans a workspace's full message history.
- Workspace agent inherits the "dialectic_chat" trace name; scope stays
distinguished by agent_type/track_name (workspace name was already in
telemetry context).
- Prefetch failure logs carry workspace + traceback; prompt no longer
contrasts against a peer-level agent the model has no concept of;
drop ticket identifiers from comments.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add `scope` to workspace chat and exclude scope peers from stats
Workspace chat is peer-unanchored, so `scope` is always a session-union
allowlist (single name or list), fail-closed when empty. Stats and
active-peer prefetch drop scope-kind peers and honor the same allowlist.
* test: teach the unified runner `workspace_chat` and parse every case
QueryAction now accepts target=workspace_chat (SDK path, including
scope). A pytest over tests/unified/test_cases/*.json keeps the four
existing workspace-chat cases — and a new scoped one — from rotting
against the schema again.
* docs: tighten workspace-chat scope docs and judge prompt
Scoped workspace_chat uses the SDK, not raw HTTP. The scope fixture's
judge now requires the in-scope tea fact, not merely the absence of the
leak. format_sse_stream matches the peer-chat one-liner.
* fix(dialectic): restore the empty-memory fallback for workspace chat
`search_memory` auto-searches messages when a pair has no observations,
but the gate only admitted `agent_type == "dialectic"`. The workspace
executor passes `workspace_dialectic`, so workspace chat got a bare
"No observations found" and answered that it knew nothing rather than
falling through to message search.
Also fixes the two unified cases that never ran: `deriver` is not a
field on `WorkspaceConfiguration`, so both aborted at load with
`extra_forbidden`. `workspace_chat_scope` additionally enables reasoning,
since it asserts scope isolation and has no reason to depend on the
fallback path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(tests/unified): fail CI when unified tests fail
`runner.run()` tallied failures into `failed_count` and printed them, but
returned nothing, and both entrypoints ignored the result. The workflow
invokes `python -m tests.unified.run` bare, so the job has gone green on
failing and unrunnable cases since it was wired up in #291.
Return the count and exit non-zero on it. `INVALID SCHEMA` already counts
toward the tally, so a malformed case now fails the job instead of being
skipped silently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(unified): assert scope peers stay out of workspace chat answers
Scope peers are real peer rows, so a regression in the `scope_peer_clause`
exclusion would surface `scope.therapy` through workspace stats or the
routing prefetch. Nothing asserted against that.
Adds the check to the existing scoped query and a new unscoped one, since
the two exercise different `get_active_peers` branches. Verified by
removing the exclusion, which fails the unscoped query.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: Remove dead code references
---------
Co-authored-by: doria <93405247+dr-frmr@users.noreply.github.com>
Co-authored-by: Benjamin McCormick <docterformer@protonmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Aakash Kattelu <aakash@plasticlabs.ai>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix(filter): make ne on jsonb metadata keys null-safe
* test(filter): cover null-safe ne on nested metadata keys
* test(filter): count actual rows for nested-metadata ne null-safety
Compile-only checks lock the operator map entry but don't catch wrong
row sets under three-valued logic. Adds a live messages/list case with
a message missing the key and one with empty metadata, following the
scalar-column pattern in test_negation_includes_conclusions_with_no_session.
* test(filter): type message_configs with a TypedDict
basedpyright couldn't narrow the heterogeneous metadata dict literals,
so indexing message_configs["content"] came back partially unknown and
broke the sorted() calls under type checking.
* fix(deriver): truncate oversize observations so one cannot drop the batch
simple_batch_embed raised ValueError when any input exceeded the per-input
token cap, which failed the entire deriver save when a single observation
was over-length. Add on_oversize="truncate": oversize inputs are embedded
from a token-capped prefix (re-encoded until it fits, with a warning),
preserving one vector per input. Default stays "raise" so existing callers
are unchanged. RepresentationManager opts into truncate.
Also add a live embedding test that fails on main (raise / missing kwarg)
and passes once a mixed short+oversize batch survives.
Refs #569
* fix(deriver): surface failure when all observer saves fail
When every observer's save_representation failed (e.g. embedding retries
exhausted under a sustained 429), the deriver logged the error and returned
normally, so the queue marked the work unit processed with zero documents
saved. Collect per-observer errors and, after telemetry is emitted, raise
RepresentationSaveError when no observer succeeded. Partial failures stay
processed (saved observers must not be discarded) and are recorded via an
additive failed_observer_count on RepresentationCompletedEvent.
Refs #728
* fix(embedding): guarantee truncation progress and truncate on re-embed
The retry slice in _truncate_to_token_limit always recomputed the same
keep count, so a slice whose re-encode grew past the cap could oscillate.
Decrement keep after each unsuccessful retry.
Document re-embed in the reconciler used the default on_oversize="raise",
so one oversize document failed every other document in the batch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: drop ticket ids and shrink comments to one sentence
Comments and docstrings describe current behavior, not the PR that
introduced them. Ticket numbers stay in the commit/PR.
* chore: annotate RepresentationSaveError and assert truncate on re-embed
* fix(embedding): truncate on conclusion create paths and document BPE loop
Storage callers in create_observations (API + agent tools) now pass
on_oversize="truncate" so a single oversize item cannot drop the batch.
Docstring on _truncate_to_token_limit notes why decode/re-encode is load-bearing.
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* telemetry: materialize dropped-event counter children at 0
A labeled Prometheus counter exports no series until its first labels()
call, so telemetry_events_dropped stayed invisible until an event was
actually dropped — impossible to alert on or graph, and "no drops" was
indistinguishable from "metric missing / scrape broken".
Pre-create the (namespace, reason) children at 0 on emitter start, for
each reason the emitter can emit, so the metric is always present.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* telemetry: generalize counter zero-init to all bounded-label counters
Extends #927 (which zero-inited telemetry_events_dropped) to every counter
whose label domain is bounded and known at startup, so metrics are present in
Prometheus before their first event — a missing series then signals a broken
scrape rather than "nothing happened yet".
- add initialize_bounded_metrics(instance_type) on PrometheusMetrics; call it
per-process from main.py (api) and deriver/__main__.py (deriver).
- extract a shared _touch() helper; refactor initialize_telemetry_dropped_metrics
onto it (that one stays per-emitter in start() — it's prefix-dependent).
- explicit ALL_EVENT_TYPES / HIGH_VOLUME_EVENT_TYPES registry in telemetry.events,
drift-guarded by tests that walk BaseEvent subclasses.
- only VALID (task_type, token_type, component) tuples for deriver_tokens (the
cartesian product would fabricate impossible always-0 series); only high-volume
event types for sampled_out; high-cardinality labels (endpoint, workspace_name)
left open.
- gauges: zero-init embed_now_tasks_in_flight + telemetry_buffer_size; add a new
message_embeddings_pending backlog gauge, set each reconciliation cycle and
zero-inited at deriver startup (Rajat's pending/in-flight ask).
- backfills the tests #927 shipped without.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* review: task-aware deriver combos + fail-soft gauge zero-init
I1: _DERIVER_TOKEN_COMBOS was factored task-independently, materializing the
impossible (ingestion, input, previous_summary) series — previous_summary is
summary-only. Make combos task-aware (_DERIVER_TOKEN_COMBOS_BY_TASK) so no
always-0 impossible series is fabricated, matching the PR's own goal. Tests
tightened to assert the ingestion/previous_summary series is absent.
I2: the three gauge .set(0) zero-inits were bare while the counter inits go
through the fail-soft _touch. Add _set_gauge_zero() so a gauge init can't
propagate an exception into process startup either.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(telemetry): isolate zero-init namespaces, add deriver-to-api guard
Global-REGISTRY assertions used a fixed "test" namespace, which several
other suites also pin, so another test's materialized children could
satisfy a presence assertion or break an absence one. Each test now runs
under a unique namespace resolved from settings at read time.
Adds the inverse per-process isolation test: deriver-only init must not
materialize API-only series (dialectic tokens, embed_now).
Co-Authored-By: Claude <noreply@anthropic.com>
* review: per-replica backlog gauge, drop duplicated constants and .meta refs
Addresses Vineeth's review on #927.
Blocking:
- message_embeddings_pending is a DB-global count, so drive it from
ReconcilerScheduler._scheduler_loop (runs on every replica, every
interval) instead of run_vector_reconciliation_cycle (runs off the
queue behind work-unit dedup, so one replica per cycle). Combined with
the zero-init, the old placement made every replica that never won the
work unit export a confident permanent 0. Help string now names the
owner so dashboards don't reach for sum().
- guard initialize_telemetry_dropped_metrics on METRICS.ENABLED,
matching its sibling initializer.
- drop the duplicate REASONING_LEVELS; import the one in src/config.
Non-blocking:
- walk BaseSpecialist recursively via a shared utils.types.walk_subclasses
(replaces the direct-children-only __subclasses__() and the test's
private copy of the same helper).
- derive the specialist assertion from the subclasses instead of
hardcoding two names — the hardcoded pair kept passing after
CardRefreshSpecialist landed, leaving it uncovered.
- inline the zero-init rationale and the multi-instance bucket taxonomy;
removes both pointers to a .meta design doc that is not in the repo.
Tests: new tests/reconciler/test_pending_backlog_gauge.py pins both
halves of the relocation (verified it fails when reverted).
* review: fix inert test guard, stale comments, and the REASONING_LEVELS drift claim
Second review pass on the branch. Findings, most severe first:
- tests/reconciler/test_pending_backlog_gauge.py: the _try_enqueue_task stub
was patched onto the class but declared without `self`, so calling it
raised TypeError — which _scheduler_loop swallows. The guard was inert and
the test passed for the wrong reason. Fixed the arity.
- metrics.py still commented that the backlog gauge is "set live each
reconciliation cycle". That is the exact claim the previous commit
overturned; it now contradicted the help string, the bucket-3 docstring
and sync_vectors.py.
- metrics.py claimed REASONING_LEVELS is "derived from the config Literal so
it never drifts", but config.py hand-listed it, so the earlier dedup had
quietly traded away the guarantee the original get_args() call provided.
Made it true instead: config.REASONING_LEVELS = list(get_args(...)), which
keeps the dedup and restores the invariant.
- dropped _set_gauge_zero: all three gauges it zeroed already have identical
fail-soft setters, so it was a second way to do one thing. Using the
setters also makes _handle_metric_error name the actual gauge.
- record_pending_embeddings_backlog's docstring oversold the covering index
as making the COUNT "negligible". The index makes cost proportional to the
pending backlog, not to the table — which is worst precisely when the
backlog matters. Stated honestly.
- _scheduler_loop's docstring said it only enqueues; it also refreshes the
gauge, at a cadence set by the shortest task interval.
- comment reconciliation: stripped #927 / "the generalization" temporal
anchoring, a CardRefreshSpecialist change-narration clause, and
reviewer-directed phrasing from the test file; disambiguated the
src/utils/summarizer.py path.
- CLAUDE.md had no Prometheus section at all, so the new "add a BaseEvent
subclass -> update ALL_EVENT_TYPES" obligation and the never-sum() rule
for non-additive gauges were undiscoverable from the architecture doc.
Verified: ruff + basedpyright clean (0 errors), tests/telemetry + reconciler
+ dialectic + llm 497 passed, full suite 1768 passed with only the 4
pre-existing test_document failures (OpenAI key required, reproduced on
clean origin/main). Re-confirmed the relocation guard fails when reverted.
* fix: silence the two basedpyright warnings inherited from main
CI runs `uv run basedpyright` bare, and basedpyright exits non-zero on any
warning — so these two have been failing the staticanalysis job on every
branch cut from current main, not just this one:
- src/vector_store/__init__.py:209 implicit string concatenation (#496)
- tests/test_cache_redaction.py:5 private import (#869)
Both predate this branch and are unrelated to the telemetry work; fixed
here only because they block this PR from going green. Verified: clean
origin/main also reports "0 errors, 2 warnings" and exits 1.
basedpyright now 0 errors, 0 warnings, exit 0.
* docs(telemetry): make the bucket-3 aggregation rule precise
The multi-instance taxonomy said a service-scoped non-additive metric has
"no aggregation correct once they disagree", then immediately mandated that
every instance refresh on its own timer. Those undercut each other: staggered
timers ALWAYS disagree slightly, so as written the rule reads as "ensure they
don't", which is unachievable, and it leaves the reader unsure whether max()
and avg() survived the fix.
The actual rule is bounded disagreement plus a scale-preserving aggregator.
Instances are N witnesses to one fact, not N parts of one whole, so sum() can
never be correct (it scales with replica count) while max()/avg()/quantiles
are correct precisely because the per-instance timer bounds the spread.
Wording only; no behavior change. The gauge help string already said
"max() or avg(), never sum()" — this makes the normative docstring agree
with it. Surfaced walking Vineeth's comment 3668208059 for comprehension.
* refactor(bench): import REASONING_LEVELS from config instead of re-listing
Third copy of the constant, missed when ee781c0/694e07f deduped the other
two. This one re-declared the ReasoningLevel Literal as well as the list,
so the type alias could diverge from config's with nothing to catch it —
and the list was hand-written, the variant that typechecks clean while
missing a member.
No import barrier justified it: this module already imports from src, as do
seven of its siblings in tests/bench. Concrete effect of the drift was that
a newly added sixth reasoning level would be rejected by the bench CLI's
argparse choices=.
src.config.REASONING_LEVELS is now the single definition repo-wide.
* test(telemetry): pin the METRICS.ENABLED guard on the per-emitter initializer
initialize_telemetry_dropped_metrics gained a METRICS.ENABLED guard in
ee781c0, addressing Vineeth's asymmetry comment, but nothing asserted it —
it had only the enabled half of the pair its sibling has. Deleting the guard
left the suite green, so the fix closed the asymmetry in the guards and
reproduced it one level up in the tests.
Mirrors test_init_noop_when_metrics_disabled. Verified live rather than
assumed: deleting the two guard lines turns this test red.
Uses a unique namespace, without which the absence assertion would be
satisfied by the enabled test's children rather than by the guard.
* docs(telemetry): fold zero-init why-prose behind # region ai markers
Comment/docstring-only pass over the changed files, per the groudon
comment-marker standard: the terse human-facing "what" stays visible, and
load-bearing "why" (the zero-init / absent-series-means-broken-scrape
rationale, gotchas, receipts) folds into # region ai / # ai: blocks.
Behavior-preserving: AST-identical modulo docstrings/comments vs the
pre-pass merge; ruff, ruff format --check, and basedpyright all clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: honor DERIVER_DEDUPLICATE in create_observations
The agent-tool path hardcoded deduplicate=True, so DERIVER_DEDUPLICATE=false
could not disable dedup for observations created through this path. Pass
settings.DERIVER.DEDUPLICATE, matching crud/representation.py.
* test: cover deduplicate setting is forwarded in create_observations
* feat: scopes SDK surface and session allowlist on session context
Exposes the Scopes v1 facade in both SDKs, which until now was reachable
only by hand-rolled HTTP, and closes the Phase 1 gap where the session
allowlist never landed on the context route.
SDKs (DEV-2001, folds in DEV-1996)
New Scope class in both SDKs — addSessions / removeSession / sessions /
status — plus honcho.scope() and honcho.scopes() entry points, a
`scopes` option on session creation, and `scope` + `sessions` read
options on chat, chatStream, representation, and session.context.
`scope` on workspace search. Python covers sync and .aio equally.
`sessions` is sugar, not a new wire field: on the recall endpoints it
goes out as the constrained `filters: {session_id: [...]}` body, never
as a key of its own. Kept separate from the `filters` parameter on the
list/search methods on purpose — that one is the full filter DSL,
whereas the recall endpoints accept a single key and 422 on anything
else, so one name for two grammars would be a trap.
Server (DEV-2357)
`GET /sessions/{id}/context` accepts a `sessions` allowlist confining
the target's representation. Two deliberate choices worth review:
- Sent as a repeated query parameter rather than the `filters` body the
issue specced. The route is a GET and `session_id` is the only
supported key, so a JSON blob in a query string buys nothing.
- The peer card is omitted under an allowlist. Cards key on
(workspace, observer, observed) with no session dimension, so they
cannot be narrowed; returning one would leak exactly what the
allowlist exists to exclude. Same reasoning as ALLOWLIST_SAFE_LEVELS.
`scope` needs no carve-out — it swaps the observer to the scope peer,
so the card read is the scope's own.
`extract_session_allowlist` now delegates to a shared
`normalize_session_allowlist`, so the cap, id charset, and must_include
rule have one implementation across both entry points. Existing error
messages are unchanged.
Also in here
- ConclusionScope renamed to ConclusionsView in both SDKs. "Scope" now
means a named set of sessions, which that class is not — it is a view
over one observer/observed pair. ConclusionScope kept as a deprecated
alias; the package-level import path only.
- The TS HTTP client comma-joined array query params, so any list-valued
parameter arrived as one malformed entry. Fixed at buildURL rather
than the call site.
Not in this PR: the "How Scopes Work" docs guide and the Groudon
dashboard tab (both DEV-2001), and CHANGELOG entries for Phases 2a-2c,
which are still merged-but-unrecorded.
Verified: ruff, basedpyright, tsc --noEmit, biome all clean. New unit
tests cover the SDK option translation and the Scope client, but the
context route's own behavior — the 422s, the 401 membership gate, the
dropped peer card — has no test yet; the analogous chat/representation
cases in tests/test_session_allowlist.py are the place for it.
Refs DEV-2001, DEV-1996, DEV-2357
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Exposes the Scopes v1 facade in both SDKs, which until now was reachable
only by hand-rolled HTTP, and closes the Phase 1 gap where the session
allowlist never landed on the context route.
SDKs (DEV-2001, folds in DEV-1996)
New Scope class in both SDKs — addSessions / removeSession / sessions /
status — plus honcho.scope() and honcho.scopes() entry points, a
`scopes` option on session creation, and `scope` + `sessions` read
options on chat, chatStream, representation, and session.context.
`scope` on workspace search. Python covers sync and .aio equally.
`sessions` is sugar, not a new wire field: on the recall endpoints it
goes out as the constrained `filters: {session_id: [...]}` body, never
as a key of its own. Kept separate from the `filters` parameter on the
list/search methods on purpose — that one is the full filter DSL,
whereas the recall endpoints accept a single key and 422 on anything
else, so one name for two grammars would be a trap.
Server (DEV-2357)
`GET /sessions/{id}/context` accepts a `sessions` allowlist confining
the target's representation. Two deliberate choices worth review:
- Sent as a repeated query parameter rather than the `filters` body the
issue specced. The route is a GET and `session_id` is the only
supported key, so a JSON blob in a query string buys nothing.
- The peer card is omitted under an allowlist. Cards key on
(workspace, observer, observed) with no session dimension, so they
cannot be narrowed; returning one would leak exactly what the
allowlist exists to exclude. Same reasoning as ALLOWLIST_SAFE_LEVELS.
`scope` needs no carve-out — it swaps the observer to the scope peer,
so the card read is the scope's own.
`extract_session_allowlist` now delegates to a shared
`normalize_session_allowlist`, so the cap, id charset, and must_include
rule have one implementation across both entry points. Existing error
messages are unchanged.
Also in here
- ConclusionScope renamed to ConclusionsView in both SDKs. "Scope" now
means a named set of sessions, which that class is not — it is a view
over one observer/observed pair. ConclusionScope kept as a deprecated
alias; the package-level import path only.
- The TS HTTP client comma-joined array query params, so any list-valued
parameter arrived as one malformed entry. Fixed at buildURL rather
than the call site
* fix(scopes): close peer-card leak under limit_to_session, harden SDK inputs
Addresses review findings on the scopes work. All four were verified by
reproducing them, not by reading.
Peer card no longer leaks under any allowlist
The card was dropped when `sessions` was set but returned when
`limit_to_session=true` produced the identical allowlist, so a control
meant to fail closed was defeated by swapping one query parameter. It is
now gated on the effective allowlist, computed once and shared by the
representation call and the card read — the duplicated inline
conditional is what let the two drift apart.
`POST /peers/{id}/chat` still injects an unscoped card under an
allowlist (src/dialectic/chat.py fetches it on peer_card.use alone, with
no reference to session_allowlist). Left alone deliberately: that is a
behavior change to the shipped dialectic and
Scope validation messages survive the option union
ScopeOptionSchema is a union, and Zod collapses a failing union into one
`invalid_union` / "Invalid input" issue, burying the branch errors. Every
invalid scope on chat/representation reported "Invalid input" and told
the caller nothing — including the reserved-prefix case the check order
exists to surface. The rules are now a plain function applied after the
union resolves, so the specific message reaches the caller for bad
charset, reserved prefix, empty and over-cap lists alike.
Empty scope no longer fails open
`session.context({scope: ''})` and `honcho.search(q, {scope: ''})` used
truthiness checks, so an invalid scope was dropped and the call returned
*unscoped* results. Both now test against undefined so the value reaches
the schema.
Session IDs validated before reaching a URL path
`scope.removeSession('valid-session?typo')` addressed `valid-session`
with a stray query string: the wrong session removed, and reconciliation
run against it. Both SDKs now validate the charset first. Python's
`add_sessions` was unvalidated too — harmless in a JSON body, but
leaving one path checked and its sibling unchecked is how this recurs.
Also
- Corrected the `limit_to_session` description: it claimed "only used if
search_query is provided", but the allowlist reaches
_query_documents_recent unconditionally.
- Corrected the documented 1,000-session cap on `sessions`, which is
unreachable via repeated query params — the request line exceeds h11's
16 KB and nginx's 8 KB defaults at a few hundred entries, giving an
opaque 414/431 instead of a 422.
- Removed a dead route builder.
Tests: 7 new TypeScript cases and 2 new Python classes covering all four
findings. The TypeScript unit suite passes 146/146. The context route
itself is still unexercised — the card gate and the 401 membership check
remain verified by reading only.
Refs DEV-2001, DEV-1996, DEV-2357, DEV-2201
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sdk): align the conclusions-view error message across both SDKs
The ConclusionScope -> ConclusionsView rename updated the identifier but
not the prose inside the thrown message, which the rename pattern
(\bConclusionScope\b) does not match. TypeScript ended up throwing
"managed by this conclusions view" while Python still threw "managed by
this conclusion scope" — the same error, different text per SDK.
Three server-backed conclusions.test.ts cases assert that message by
regex and failed under `pytest -k typescript`. The four equivalent Python
assertions were passing, because they matched Python's unchanged string —
so fixing only the TypeScript tests would have made the suite green with
the divergence still in place.
Brings Python's message, comments and docstring in line with TypeScript,
and updates the assertions in both suites. `grep -ri 'conclusion scope'`
is now empty.
Verified: 64 passed across tests/sdk_typescript/, tests/sdk/test_conclusions.py
and tests/sdk/test_scope_options.py — the last of which had never been run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: enforce explicit-document session purity in dedup/merge paths
Audit for DEV-2000 (Scopes RFC prerequisite): explicit-level documents must
stay session-pure so scope memory can be built by copying explicit documents
between collections. Two classes of violation were possible:
- Exact-content and semantic dedup in crud/document.py matched candidates
with no level or session scoping, so an explicit document could be
reinforced by — or soft-deleted in favor of — a same-content document from
a different session or a different level (silently merging cross-session
derivations into one row).
- The generic create_observations tool handler accepted level='explicit'
from agents with no message context (dreamer/dialectic), which would mint
session-less explicit documents.
Enforcement (refuse, never rewrite):
- create_documents refuses explicit documents with a null session_name
- exact dedup keys on (content, level, session-for-explicit); derived levels
keep cross-session consolidation
- is_rejected_duplicate scopes candidate search to the same level, and the
same session for explicit documents
- the create_observations tool rejects explicit-level input outside message
ingestion (deriver) context
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add card_refresh dream type for event-driven peer-card updates
Adds a lightweight dream variant (DEV-2000, Scopes RFC prerequisite) that
runs ONLY the peer-card update — for event-driven refreshes such as scope
membership changes and cold starts:
- DreamType.CARD_REFRESH alongside OMNI; dispatched by process_dream to a
new run_card_refresh_dream orchestration
- CardRefreshSpecialist: restricted to get_recent_observations,
search_memory, and update_peer_card (no observation-mutating tools), with
a low tool-iteration cap of min(6, DREAM.MAX_TOOL_ITERATIONS)
- rebuild=True mode carried in the dream payload: the existing card is NOT
injected into the prompt and the specialist rebuilds it solely from
observations present in the collection (for use after removals)
- enqueue-able via the manual enqueue_dream path (bypasses volume gates);
the work-unit key already embeds the dream type so a card refresh never
collides with a pending omni dream. POST /v3/workspaces/{id}/schedule_dream
accepts dream_type=card_refresh plus the rebuild flag
- card refreshes never advance the omni dream guard pair
(last_dream_at / last_dream_document_count)
- shared PEER CARD prompt section extracted (verbatim) from
DeductionSpecialist for reuse; CallPurpose gains dream.card_refresh
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: reserve scope__ peer namespace with kind flag and guardrails
Introduce the scope peer namespace (scope__<name>) and the authoritative
{"kind": "scope"} configuration flag, plus the server-side guardrails:
- src/utils/scopes.py: single source of truth for the prefix, kind flag,
and name helpers (scope_peer_name / is_scope_peer_name /
scope_name_from_peer / validate_no_scope_peer_names)
- reject reserved-prefix names on peer get-or-create (422)
- reject scope peers as message authors in crud.create_messages (422)
- reject scope peers as chat/representation targets (422); a scope peer
as the path-level observer is deferred to Phase 2b
- reject scope peers on the generic session-peer add/set/remove routes
and the session-create peers mapping (422, directing to scopes routes)
- peers.list excludes scope peers by default; new PeerGet.kind option
("scope" | "all") switches the view via a configuration JSONB filter
- schemas: Scope / ScopeCreate / ScopeSessions(Add) and
SessionCreate.scopes (unprefixed scope names, validated)
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add scopes CRUD routes and session-create scopes wiring
New /v3/workspaces/{workspace_id}/scopes facade (workspace-level auth;
peer- and session-scoped keys are rejected):
POST "" create-or-get (201/200)
POST /list paginated scope list
GET /{scope_id} single scope
POST /{scope_id}/sessions add memberships
DELETE /{scope_id}/sessions/{session_id} remove membership
GET /{scope_id}/sessions list member session ids
- crud/scope.py: get_or_create_scopes stamps the backing peer with
{"kind": "scope", "observe_me": false} and refuses to adopt a
legacy peer occupying the reserved name without the flag (409)
- memberships are session_peers rows with observe_others=true /
observe_me=false — identical to a hand-built observer peer
- SessionCreate.scopes: create-or-get each scope peer and add the
membership at session creation (the no-backfill common path)
- crud/session.py: public upsert_session_peers wrapper so the facade
bypasses the route-level guardrails without reaching into privates
Backfill of pre-existing documents and reconciliation on removal land in
DEV-1999; membership only affects messages ingested after the change.
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: cover scopes facade, guardrails, and observer semantics
- create-or-get idempotency, list/get, name validation, legacy-collision
rejection (409), auth scoping (workspace key ok, peer/session keys 401)
- reserved prefix rejected on peer create; peers.list kind filtering
- scope peers rejected as message authors, chat/representation targets,
and on the generic session-peer routes
- membership add/list/remove with observe_others=true / observe_me=false
row shape asserted via DB, and facade-less equivalence with a
hand-built observer peer
- end-to-end litmus: after adding a session to a scope, the deriver
enqueue fan-out includes the scope peer as an observer
- session creation with scopes: [a, b] creates both memberships
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: scope backfill-by-copy and removal reconciliation jobs
Retroactive scope membership changes (DEV-1999):
- New queue task types scope_backfill / scope_removal with payloads,
work-unit keys ({task}:{workspace}:{scope_peer}:{session}), deduped
enqueue (mirrors enqueue_dream), and consumer dispatch.
- Backfill copies a session's explicit documents from each sender's
global (P, P) collection into the scope's (scope_peer, P) collection
with internal_metadata.copied_from as the idempotency marker;
soft-deleted copies from an earlier removal are restored, so
add -> remove -> re-add converges on exactly one live copy. Completion
enqueues one manual omni dream per touched collection.
- Removal soft-deletes the session's explicit documents in the scope's
collections and cascades (fail-closed, transitively) to derived
documents whose source_ids intersect anything removed, deletes the
vectors from the external store, then enqueues a card_refresh dream
with rebuild=True plus a manual omni dream per touched collection.
- Zero LLM re-derivation: explicit documents are session-pure (DEV-2000
invariant); the only external call is re-embedding rows whose
embedding column is NULL (external-store deployments).
- Per-session job status lives in the scope peer's internal_metadata
under backfill_status, written via single-statement JSONB merges
(concurrent-writer safe) and surfaced at
GET /v3/workspaces/{w}/scopes/{scope_id}/status.
- Enqueued from the scopes add-sessions route and SessionCreate.scopes
handling, only when the session already has messages.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: add scope backfill/removal test coverage, fix backfill status JSONB bug
Adds tests/deriver/test_scope_backfill.py covering the DEV-1999 scope
backfill-by-copy and removal reconciliation jobs implemented in a prior
commit: explicit-doc copying, copied_from idempotency (including
add->remove->re-add), multi-peer collection routing, removal cascade to
dependent derived docs, dream enqueues (manual omni on backfill;
card_refresh rebuild + omni on removal), the status route, and add-sessions
route wiring (backfill enqueued only when the session already has messages).
Fixes a real production bug surfaced by these tests: update_scope_backfill_status
passed json.dumps()'d strings through SQLAlchemy cast(..., JSONB), which
double-encodes (psycopg re-serializes the already-JSON string), producing a
JSONB string scalar instead of an object. Postgres's `||` between two
non-array jsonb scalars doesn't merge — it silently wraps both into a
2-element array, corrupting backfill_status into a list. This crashed
clear_scope_backfill_status's `#-` path delete (called on every removal)
with "path element is not an integer" once a session had ever completed a
backfill. Fixed by passing raw Python dicts to cast() instead, mirroring the
working pattern already used in update_collection_internal_metadata.
Also adds src.deriver.scope_backfill.tracked_db to conftest's tracked_db
patch list — the module was missing from that per-import-site allowlist, so
its DB work ran against the real configured database instead of the
isolated per-test database.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(crud): preserve cache invalidation across get_or_create retry
`get_or_create_peers` and `get_or_create_scopes` mutate existing rows, then
insert new ones inside `db.begin_nested()`. When a concurrent writer creates
one of those rows first, the insert raises IntegrityError and the function
retries.
`begin_nested()` autoflushes the pending UPDATEs *before* opening the
savepoint, so the rollback neither undoes them nor expires the now-clean ORM
state. The retry then compared already-updated values, found no change, and
dropped those peers from `changed_peers` — skipping the cache purge while the
row change committed anyway, leaving entries stale until the 300s TTL.
Carry the mutated names into the retry via `_pending_invalidation` so the
purge cannot be lost.
The scopes facade mirrors `get_or_create_peers`, so both copies carried this.
The peer path is pre-existing and runs on every message ingest.
Also add the missing /v3 prefix to `_SCOPES_ROUTE_GUIDANCE`, which pointed
callers at a 404.
Adds tests/crud/test_get_or_create_retry_invalidation.py, which drives a real
racing session and fails without this change.
* fix(scopes): make scope identity unforgeable, unblock non-pattern peer names
Three coupled changes to the scopes facade.
1. `PeerCreate` no longer gates internal lookups. It exists to validate a new,
user-supplied peer id at the API boundary, but crud used it as a DTO for
names that already exist, so any name outside RESOURCE_NAME_PATTERN raised a
raw pydantic ValidationError — which is not a HonchoException, so it fell
through to the catch-all handler as an HTTP 500. Adds `PeerSpec` (same
fields, no charset pattern) as `PeerCreate`'s base, widens
`get_or_create_peers` to accept it, and changes `get_peer` to take a plain
str. All 13 construction sites converted; the create route keeps full
validation.
This unbreaks the Dreamer: DreamScheduler passes `collection.observer`
straight into the specialist preflight, and scope peers have
`observe_others=true`, so every `(scope.x, peer)` dream died there — the
feature scopes exist to enable. It also fixes a pre-existing bug unrelated
to scopes: a peer named `alice.smith` (legal before d429de0e5338, which
validated names by length alone) 500s on message create, session peer add,
and peer update.
2. The `kind` flag moves from `configuration` to `internal_metadata`.
`configuration` is user-writable — `PeerCreate`/`PeerUpdate` accept a
free-form dict and `update_peer` replaces it wholesale — so a legitimate
`{"observe_me": true}` update silently dropped the flag, and a forged
`{"kind": "scope"}` injected an ordinary peer into `POST /scopes/list`.
`internal_metadata` appears in no API schema. `observe_me: false` stays in
`configuration`, where it belongs.
3. Scope identity requires prefix AND flag, via `is_scope_peer()` and
`scope_peer_clause()`. Neither half is forgeable: the prefix sits outside
RESOURCE_NAME_PATTERN, `internal_metadata` is unreachable. Usage-site guards
become flag-based so a legacy peer merely occupying the namespace keeps
working rather than 422-ing on its own traffic; peer create and update stay
name-based, since those must stop new names entering the namespace.
`update_peer` now returns 422 instead of 500.
Also swaps the reserved prefix from `scope__` to `scope.`: `_` is inside
RESOURCE_NAME_PATTERN, so any tenant could already own a `scope__x` peer.
No DB migration — `internal_metadata` already exists on `peers`, and no scope
peers exist in any deployment yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): validate peer names on create, close namespace squatting and upsert race
Addresses three review findings against 48047a6a.
1. `PeerSpec` let API callers create invalid and reserved-prefix peers.
Widening `get_or_create_peers` to accept a pattern-free schema fixed the
lookup 500s but also removed validation from the *insert* path, and
request-controlled names reach it via message authors, session peer maps, and
the chat observer path — none of which carry a charset pattern of their own.
Confirmed: `POST /sessions/{id}/messages` with `peer_id: "scope.x"` returned
201 and minted an unflagged squatter, after which `POST /scopes {id: x}` was
permanently 409-blocked — namespace denial of service by any caller able to
post a message. `peer_id: "not a valid name!@#"` was likewise created.
Fixed by validating only names about to be INSERTed
(`_validate_new_peer_names`), so already-existing names — legacy dotted
names, scope peers — still resolve without a spurious 422. That keeps the
Dreamer fix intact, since it reads through `get_peer`.
2. Existing reserved-prefix squatters could not be updated. The name-based guard
on `PUT /peers/{peer_id}` refused every `scope.` name, contradicting the
invariant that an unflagged squatter stays a normal peer. Now flag-based, so
behavior is three-way: a real scope is refused, an existing unflagged peer
updates, and a missing reserved-prefix name is refused by (1) rather than
minted.
3. Scope checks raced with get-or-create and the membership upsert. The
route-level guards run before peers are resolved, so a scope created
concurrently in that window would be attached by the generic path with a
default `SessionPeerConfig()`, clobbering its observer membership config.
Adds `_reject_resolved_scope_peers`, which runs on the resolved rows in the
same transaction as the upsert — no window, no extra query. The early checks
stay for better error messages.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): guard scope membership config, move checks to the mutation point
Addresses a second review pass against 10655792.
1. Scope membership configuration was directly user-mutable.
`PUT /sessions/{id}/peers/{peer_id}/config` had no scope guard at all, and
`crud.set_peer_config` resolved the peer only to discard the row. Confirmed:
posting `{"observe_others": false, "observe_me": true}` for a scope returned
204 and persisted, which silently stops all fan-out into the scope and makes
Honcho form a representation *of* a scope — neither of which is reachable
through the facade. Deterministic, no race required. Now checked on the row
`get_peer` already returns, so it costs nothing and cannot race.
2. Empty and over-long names were still 500s. Removing the charset pattern from
`PeerSpec` fixed one trap but left its length bounds, and request-bound peer
names carry no length limits of their own — so `peer_id: ""` or a 513-char
name reached `PeerSpec(...)` and raised a raw pydantic ValidationError that
the catch-all turned into a 500. `PeerSpec` now carries no constraints at all
(matching its documented purpose) and every rule for a new name lives in
`_validate_new_peer_names` on the insert path.
3. Resolved-row protection generalized. The previous pass applied it only to
membership upserts, leaving check-then-use windows elsewhere: peer update
could have a concurrently-created scope's configuration replaced wholesale
(create-path validation does not fire for a peer that now exists), the chat
observer get-or-create could resolve a fresh scope as its observer, and the
generic session-peer removal could silently detach a scope from its sessions.
Each now inspects the resolved peer immediately before acting; the redundant
name-level guard on the update route is dropped in favor of the race-free one.
`remove_peers_from_session` grows an internal `_allow_scope_peers` flag because
the scopes facade ends membership through that same path and must not be blocked
by its own guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(scopes): enumerate every peer-touching route against a scope policy
Three review passes each found the same class of defect: a route nobody had
checked, rather than logic that was subtly wrong. One of them — `PUT
/sessions/{id}/peers/{id}/config`, which let any caller set a scope to
`observe_others=false` and silently stop all fan-out into it — predated this work
entirely, because the guardrail set was assembled guardrail-by-guardrail instead
of derived from the route list. Sampling review cannot close that kind of gap;
enumeration can.
Derives every route through which a peer name can reach the system and requires
each to be classified as GUARDED or EXEMPT-with-a-reason, so a newly added
peer-touching route fails the suite until someone classifies it. Detection is the
union of path shape and a walk of the dependant tree (including sub-dependency
`Form(...)` params and nested request-body models), because neither signal alone
suffices: parameter names miss `POST /sessions/{id}/peers`, whose peer names are
dict keys, and path shape misses `messages/upload`, whose `peer_id` arrives as a
form field behind a parser dependency.
Both invariants are then asserted behaviorally, by calling the routes rather than
inspecting annotations — the guards deliberately live in crud, which is what makes
`messages/upload` guarded for free via `crud.create_messages`:
- a real scope is refused on all 11 guarded routes, and the rejection must name
the scope, so an unrelated 422 (a malformed body) cannot pass the assertion;
- an *unflagged* peer merely occupying the reserved namespace is unaffected. That
half regressed once already when `update_peer` used a name-based check.
Mutation-tested all three failure modes: disabling the `set_peer_config` guard
fails the guarded test naming that route; regressing `update_peer` to name-based
fails the squatter test; adding an unclassified peer route fails the enumeration.
Covers the HTTP surface only. Peer names also reach the system through the
deriver, dreamer, and queue, which have no route table to enumerate — noted in
the module docstring rather than implied to be covered.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): refuse a scope in every observed position; enumerate per position
Addresses a fourth review pass against 62681e4d. The headline finding is that the
previous commit's enumeration test had the wrong *model*, not a missing entry.
1. Manual conclusions could create knowledge about a scope. `POST /conclusions`
validated only that observer_id and observed_id exist, so a scope as
`observed_id` persisted a conclusion about a peer carrying observe_me=false and
created an (observer, scope) collection for it. Confirmed: 201, and it read
back. `POST /schedule_dream` had the same hole via `observed`.
The fix is positional, because the invariant is:
A scope may be an OBSERVER. A scope may never be OBSERVED.
A scope as `observer_id` is how scoped conclusions are stored and must keep
working (verified still 201); as `observed_id` it is now refused. Same split
applied to schedule_dream `observed`, the peer-card `target` (which also covers
a scope's self-card, since target-omitted collapses observed to peer_id), and
session-context `peer_target`.
2. Chat target and both representation roles kept check-to-use races. Only the
chat path-level observer was re-checked on its resolved row; the target was
checked by name and then resolved without inspecting scope identity. Both are
now checked at the dialectic preflight, where observer and observed are already
resolved — an absent name has already failed by then, and an existing squatter
cannot retroactively become a scope.
3. Generic membership removal was still racy. The adjacent SELECT narrowed the
window but could not close it under READ COMMITTED. The UPDATE now carries its
own correlated NOT EXISTS against scope_peer_clause(), so Postgres evaluates
the exclusion as part of the statement and a scope committed after the advisory
check still cannot be detached.
4. New-name validation ran after the name reached Postgres. A NUL byte passed the
request schemas and PeerSpec, then raised psycopg.DataError inside the lookup —
a 500. Values that cannot correspond to a stored row by construction (NUL
bytes, over-length names) are now refused before the query.
(Over-length names already returned 422; only the wasted query was real there.)
The enumeration test is rekeyed from (method, path) to (method, path, position).
A binary per-route verdict cannot express finding 1 at all: `POST /conclusions` is
one route with two positions and opposite verdicts. Detection widens to observer /
observed / target / peer_target / peer_perspective, which surfaced four routes the
previous version never saw — conclusions, schedule_dream, queue/status, and
session context.
Also registers `src.routers.workspaces.tracked_db` in the conftest patch list; the
new guard there would otherwise have run against the real configured database
instead of the per-test one.
Mutation-tested: disabling the conclusions observed-guard fails the positional
test naming that position; adding an unclassified `observed_id` param fails
enumeration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): refuse future scopes in observed positions, preserve scope membership
Addresses a fifth review pass against 14136e5b. All six findings reproduced
locally before fixing.
1. High — generic peer replacement removed scope memberships.
`set_peers_for_session` soft-deleted every active SessionPeer row, and the
request-level guard only inspected names *present* in the replacement map. A
caller detached a scope by simply omitting it, never naming it — so no
request-level guard could ever see it. Reproduced: scope sessions went
['<id>'] -> [] on a 200. The exclusion now lives in the UPDATE itself
(correlated NOT EXISTS against scope_peer_clause), so replacement means
"replace ordinary peers" regardless of request contents or concurrent creation.
2. High — peer cards could be pre-seeded for future scopes.
`set_peer_card` resolves only the observer and writes a JSONB key derived from
an unchecked observed name, and the route guard rejected only *existing*
flagged scopes. Reproduced: PUT card with target=scope.<missing> returned 200,
creating that scope then returned 201, and the card described the real scope.
3. High — dreams could be queued for future scopes.
The route checked `observed` in a read-only session that closed before
`enqueue_dream`, and a missing reserved name passes any is-it-a-scope check.
Reproduced: 204 with observed=scope.<missing>.
2 and 3 share a root cause, so they share a fix: a new `reject_scope_observed`
that is stricter than `reject_scope_peers` in exactly one case — a *missing*
reserved name is refused, because nothing on these paths creates the peer, so
nothing else would ever catch it. Existing unflagged squatters still pass.
Both guards moved to the mutation point: card validation into
`crud.set_peer_card` (same transaction as the JSONB write, so Dreamer and
agent-tool callers are covered), dream validation into `enqueue_dream` (same
transaction as the queue insert). The redundant route-level checks are dropped
rather than left as weaker duplicates.
4. Medium — prefixed NUL names still reached PostgreSQL.
`reject_scope_peers` filtered for the reserved prefix and sent matches to a
text comparison, so "scope.future\0name" raised psycopg.DataError — a 500.
Both guards now share `_reserved_name_candidates`, which materializes the input
once and rejects impossible values before any SQL. Materializing matters
independently: the message-author path passes a generator, and validation
iterates separately from the prefix filter, so a generator would be
half-consumed. `_reject_impossible_peer_names` now takes a Collection so the
type checker enforces that.
5. Medium — representation kept a check-to-use race.
The previous commit claimed both representation roles were rechecked after
resolution; that was wrong — only the dialectic preflight got that check, and
the representation route never goes through it. It now opens one short
read-only session *after* the embedding call, checks both positions, and passes
that same session to `get_working_representation`, so no connection is held
across external work and a scope committed later cannot have conclusions in the
collection being read.
6. Low — policy coverage was not exhaustive. `sender_id` reaches CRUD as
`observed` but was missing from the detected parameter set. ALLOW cases could
also not carry builders, so the suite never proved the other half of the
contract — that legitimate scope *observers* keep working, which a guard
rejecting scopes everywhere would satisfy. Both fixed; observer positions on
conclusions, dreams, cards, session context and queue status are now asserted
behaviorally.
Deliberately not implemented: the scope-creation backstop scanning for
pre-existing card keys and queue items naming a future backing peer. Reasoning is
recorded in `get_or_create_scopes` — no new such state can be created now, any
pre-existing row is coincidental since `scope.` was never a meaningful namespace,
the consequence is inert, and detecting card keys means a full table scan per
scope creation.
Mutation-tested each new guard: removing the replacement exclusion fails both
membership-preservation tests; weakening either observed guard to existing-only
fails the pre-seeding tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): exclude scope memberships from the session observer limit
Scope memberships carry observe_others=true, so every scope counted against
SESSION_OBSERVERS_LIMIT (default 10) — capping scopes-per-session at the limit
minus the session's real observers, and reporting the failure as
`400 Cannot create session <name> with 11 observers. ... Observers are peers
with 'observe_others' set to true.` on a membership call. Wrong on three counts:
the ceiling is undocumented and contradicts RFC §5.1 ("sessions belong to any
number of scopes"), the message describes session creation, and it leaks the
word "observer" through a facade whose entire job is hiding observers (RFC
goal 5). The limit exists to bound per-observer deriver fan-out for real peers;
a scope costs document rows, not LLM calls (RFC §5.2), so it does not belong in
that budget.
Excluded from both halves of the check in `_get_or_add_peers_to_session`: the
incoming names via a flag-based lookup, existing memberships via a correlated
NOT EXISTS on `scope_peer_clause()` — the same pattern the replacement and
removal paths already use, so the exclusion holds regardless of concurrent
scope creation. The early `count_observers_in_config(session.peer_names)` check
in `get_or_create_session` is left alone: `peer_names` cannot contain a scope,
and `scopes` is a separate field.
`reject_scope_peers` is split into a `scope_peer_names()` query helper plus a
two-line raiser so the observer count reuses the authoritative name-AND-flag
predicate instead of growing a third copy of it. Still costs nothing on the
common path — no reserved-prefix name in the input means no query at all.
Also caps `SessionCreate.scopes` at 100, matching `ScopeSessionsAdd.session_ids`.
This belongs in the same commit: the observer limit was the only thing bounding
that list, so removing it turns an unbounded `scopes` array into a peer row and
a membership row per element, committed — the single-request path to the
cardinality anti-pattern RFC §8 warns about. Partly answers OQ6: no per-session
cap, 100 per request.
Tests: a session joins SESSION_OBSERVERS_LIMIT + 2 scopes through both the
facade and session creation; real observers over the limit still 400, so the
carve-out cannot quietly disable the limit; 101 scopes is a 422.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): skip semantic retrieval when the embedding precompute failed
Addresses three open review comments.
1. Major — the representation read could embed inside its DB session. The
route's precompute is suppressed, and both
`RepresentationManager.get_working_representation` and
`crud.query_documents` fall back to embedding when a query arrives without
one, so a failed precompute meant an external call inside the read session
this branch opens for the scope re-check — the connection-holding rule the
route's own comment claimed to satisfy. The innermost fallback also only
catches ValueError, so a provider outage surfaced as a 500. The semantic
query is now passed only when an embedding exists, degrading to
derived+recent retrieval. (`crud.query_documents` embedding inside a caller's
session predates this branch and is left alone.)
2. Minor — `test_resolved_scope_peer_rejected_at_membership_upsert` described a
race it does not perform. It creates an already-flagged scope and calls crud
directly; the unflagged → flagged transition is not simulated. Docstring now
says what the test actually pins.
3. Minor — `test_empty_replacement_preserves_scope_membership` asserted only
half its docstring. It passed if the empty PUT left every ordinary
membership intact; now asserts the ordinary peer's left_at is set.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): close auth and observed-position gaps, paginate membership
Review response for #884.
Security:
- gate `SessionCreate.scopes` behind a workspace-level key; the session-create
route is self-authorizing, so a peer- or session-scoped token could mint scope
peers and join sessions to scopes it had no access to via `POST /scopes`
- refuse a reserved-but-nonexistent name in two observed positions that used the
permissive guard: chat `target` and session-context `peer_target`. Both let a
caller act on `scope.X` before it existed, then create the scope
Facade:
- exclude scope peers from `GET /sessions/{id}/peers` and refuse the membership
-config read for a real scope, matching its write side
- replace `GET /scopes/{id}/sessions` with `POST /scopes/{id}/sessions/list`
returning `Page[Session]`; the add route now returns 204. Membership was
unbounded on both, while every other list surface paginates
- rename `crud.get_scope` to `get_scope_or_raise`
Tests:
- add a missing-name axis to the route-policy table (`Case.refuse_missing`), which
is what surfaced the two guard gaps above
- delete 14 hand-written tests the table now enumerates; 52 -> 39 functions in
test_scopes.py with more cases covered
- tighten the squatter assertion from `!= 422` to `< 400`, which was passing on 5xx
- assert the FastAPI-internals traversal still derives positions, so a framework
upgrade can't silently empty the suite
Docs:
- drop internal ticket and RFC references from the published OpenAPI descriptions
and surrounding comments; state the behavior instead
- move implementation reasoning out of the `PUT /peers/{id}` docstring, which
FastAPI publishes, into a comment
* test(scopes): assert exact statuses for permissive missing-name cases
Follow-up review pass on #884.
- add `Case.missing_status` so a permissive missing-name position asserts the
status it should actually get (404, or 200 for the no-op removal) instead of
`!= 422`, which also passed on a 5xx — the same hole already closed in the
squatter assertion
- require it whenever `refuse_missing` is False, and require its absence when
True, so the policy table can't drift from the assertion
- repoint a stale allow-reason at POST /scopes/{scope_id}/sessions/list; the GET
it named was removed
- document the membership list's ordering under `reverse`
* chore: clean up stale docstring language
* fix: don't backfill a session that left the scope
scope_backfill and scope_removal carry different work-unit keys, so
nothing orders them: a removal enqueued right after the add — or one
that lands while the backfill is embedding — sweeps the scope before
the copies exist, leaving a departed session's documents live in the
scope forever.
_run_backfill now re-checks SessionPeer membership inside the write
transaction and returns None; process_scope_backfill then skips both
the dream enqueues and the status write, so a skipped backfill can't
resurrect the status entry removal just cleared.
Adds coverage for the skip, the NULL-embedding re-embed path, and the
failed-status write. Handler-driven tests now stand up the membership
row the guard requires.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(filter): reject unknown operator dicts on scalar columns
An unrecognized operator dict on a non-JSONB column (e.g.
{"session_id": {"operator": "null"}}) fell through to `column == value`,
binding a dict to a VARCHAR parameter. That compiles, then fails in the
driver at execute time with "cannot adapt type 'dict'" — an unhandled
500 for what is invalid input.
Raise FilterError (422) instead. The guard lives in the shared
_build_field_condition, so every route through apply_filter is covered.
It keys on the actual column type rather than the JSONB_COLUMNS name
list, so dict equality still works on JSONB columns reachable through
Document's raw-key fallback (e.g. source_ids), where the driver adapts
dicts fine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(filter): name psycopg explicitly in comment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): handle null operands and non-numeric columns in comparisons
Two defects in _build_comparison_conditions, both reachable from any
route that accepts filters:
1. A null operand hit float(None), raising TypeError where only
ValueError was caught — an unhandled 500. A null operand is a null
check, not a value comparison, so {"ne": null} now compiles to
IS NOT NULL and the other operators reject null with a 422. Equality
against null already produced IS NULL via _build_field_condition.
2. Numeric operators float()-cast on every column type, so a string
inequality on a text column ({"session_id": {"ne": "abc"}}) was
rejected as an invalid number. Coercion is now gated on the column
actually being numeric; text columns compare as text. Numeric columns
still validate, and TypeError is caught alongside ValueError.
Existing ne coverage only exercised the JSONB metadata path, which uses
_safe_numeric_cast and handles strings — the scalar column path was
untested. Adds cases for both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): fail closed on unrecognized filter shapes
The filter body is arbitrary client JSON with no schema, so validation
was emergent: any shape the DSL didn't recognize surfaced as an
unhandled 500 from somewhere in SQLAlchemy or psycopg. Fixing individual
shapes doesn't converge — a fuzz over the DSL found five more families
beyond the three already fixed here:
{"AND": [None]} TypeError, non-dict in a logical list
{"AND": [[]]} AttributeError on .items()
{"session_id": {"gte": true}} SQLAlchemy ArgumentError
{"embedding": []} NotImplementedError, no python_type
{"session_id": {"ne": {...}}} execute-time "cannot adapt type 'dict'"
Two generic guards instead:
1. Any operand bound to a non-JSONB column must be a scalar, checked
element-wise for `in`. A dict or list bound to a scalar column
compiles cleanly and only fails in the driver at execute time, so it
has to be rejected during construction. JSONB columns are exempt —
a dict there is a containment match.
2. apply_filter fails closed: FilterError propagates, anything else is
logged with logger.exception (filter shape included) and re-raised as
FilterError. Unknown filter failures become 422s while staying fully
visible as errors rather than being swallowed.
Adds two invariant tests over a generated matrix of filter shapes: every
shape either compiles or raises FilterError, and no non-scalar is ever
bound to a scalar column. Both fail without the guards above. They cover
shapes nobody enumerated, so the next unimagined body fails in CI rather
than in production.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: clear the two remaining basedpyright warnings
`uv run basedpyright tests/ src/` reported two warnings in files that
predate this change. The pre-commit hook is file-scoped, so neither was
visible unless the owning file was touched.
- src/vector_store/__init__.py: join the lancedb error message with
explicit `+` instead of adjacent literals (reportImplicitStringConcatenation).
- tests/test_cache_redaction.py: the test covers a private helper
deliberately, so annotate the import (reportPrivateUsage).
No behavior change; whole-tree check is now clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): keep numeric operands exact instead of coercing to float
float() rounds any integer past 2**53 and flattens a Decimal, so
{"token_count": {"gt": 9007199254740993}} silently compared against
9007199254740992 — a different row set than the client asked for.
_coerce_numeric passes already-numeric operands through untouched and
only parses strings, trying int() before float() so "5" stays exact
while "5.5" still parses. bool narrows to int: it is an int subclass,
but binding it as a boolean against a numeric column produces SQL
Postgres has no operator for.
Not coerced to the column's own type: int(5.5) would turn
{"token_count": {"lt": 5.5}} into `lt 5`, changing which rows match.
Also fixes a vacuous assertion in test_dict_on_jsonb_column_still_works.
It checked for "internal_metadata" in the whole statement, but that name
is in the SELECT projection either way, so the test passed even when no
WHERE clause was applied. Now asserts on stmt.whereclause and that the
filter payload is actually bound.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): bind boolean columns as booleans, numerics without a cast
Boolean columns were treated as numeric because bool subclasses int, so
`{"is_active": {"ne": true}}` coerced true to 1 and Postgres rejected
`boolean <> integer` at execute time. Confirmed against a live database:
every form except bare equality with a native boolean was a 500.
Boolean columns now get their own branch: native true/false bind as
booleans, and any other operand is a FilterError. SQLAlchemy types the
bind from the operand rather than the column, so "true" renders
`is_active = %(param)s::VARCHAR` and Postgres has no such operator — a
422 is the honest answer. String booleans have never worked, are absent
from the docs (every documented boolean is inside metadata, which is
JSONB containment and unaffected), and produced no Sentry events in 90
days, so nothing can depend on the current behavior.
Also corrects the previous commit. Coercing operands to exact ints made
SQLAlchemy render an ::INTEGER cast, so any value past int4 — not 2**53
— started failing with "integer out of range" where float() had silently
compared as a double. Decimal keeps the value exact and renders no cast,
matching what float() did. The `in` branch never went through coercion
at all, so {"token_count": {"in": [1, 2147483648]}} was a 500 before
this PR too; it now takes the same path.
Verified end to end against the live database, not just at compile time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): coerce every operand against its column's type in one place
The DSL had two operand paths with different rules. Comparison operators
parsed datetimes and coerced numbers; bare equality bound whatever it was
handed. SQLAlchemy types a bind from the operand rather than the column
and the psycopg dialect renders that type as an explicit cast, so a
mismatch compiled into valid-looking SQL and failed at execute time:
"operator does not exist: timestamp with time zone = character varying".
A matrix of column type x operand type x operator against a live
database found 462 combinations, of which 55 built cleanly and then
failed. The most plausible was a filter someone would write first try:
{"created_at": "2026-01-01"} 500
{"created_at": {"gte": "2026-01-01"}} worked
_coerce_operand now handles every operand, whatever the operator, keyed
on the column's real type: JSONB takes an object, boolean takes only
true/false, datetime parses strings, numeric goes through _coerce_numeric,
text requires a string, and a column with no python_type (pgvector) is
not filterable. eq/ne/gt/in cannot drift apart because they share the
one call; `in` coerces element-wise, since a single element's type
decides the cast rendered for that parameter. The matrix is now clean.
This is a net deletion: the separate datetime, numeric, in-datetime and
boolean branches, plus _require_bindable_operand, all collapse into it.
Two more execute-time failures fixed on the way. `contains` was keyed on
column_name == "h_metadata", so Document's equally-JSONB
internal_metadata fell through to ILIKE and produced `jsonb ~~* text`;
it now keys on the column type. And {"source_ids": "abc"} was
`jsonb = character varying`.
Closed-set columns are validated against the Literal that defines them,
so declaring a new level or sync state updates filter validation with no
change here. {"level": "banana"} was silently matching nothing.
Empty IN is now always applied rather than skipped. Unifying the branches
inherited a guard that had only ever wrapped the datetime path, which
dropped the condition entirely and widened the query to every row —
fail-open on an empty allowlist, which session scoping relies on to fail
closed (see extract_session_allowlist). Caught by an existing test that
asserts returned rows; the fuzz and the type matrix only check for
errors, so neither would have seen it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(filter): make NOT and ne include rows where the field is unset
NOT (col = v) and col <> v are NULL when col is NULL, so negation
dropped rows whose column is unset — a conclusion with no session is
not "some other session", but was excluded anyway. IS NOT TRUE and
IS DISTINCT FROM behave identically when no NULL is involved.
Adds a row-count test, since this class of bug builds and executes
cleanly, and documents {"ne": null} for excluding unset fields.
* fix(filter): reject session ids that can't name a real session
extract_session_allowlist accepted any non-empty string, so "*" reached
the three consumers of the allowlist — direct IN, the filter DSL, and a
Python membership test — which disagree about it. The DSL reads "*" as
"drop the condition" and matches every session; the others treat it as a
literal name and match none. One /chat request could have some recall
sources unscoped and others scoped to nothing.
Entries are now validated against RESOURCE_NAME_PATTERN, the same pattern
the API requires of session ids, so no session could ever be named "*"
anyway. Wildcards were never part of this endpoint's documented contract
(an id, a list of ids, or {"in": [...]}), and a wildcard alongside a
top-level session_id already 422'd via must_include.
* fix(filter): treat a bare null operand as a null check
Routing every operand through _coerce_operand made bare `None` a type
error rather than a null check, so {"session_id": null} raised FilterError
where it previously built IS NULL: _build_field_condition used to end in
`column == value`, which SQLAlchemy renders as IS NULL. Confirmed 422 on
all five column families (text, numeric, boolean, datetime, JSONB).
Nothing caught it. The docs added in this branch promise
`{"session_id": None}` matches unset rows, the comment in
_build_comparison_conditions claimed the equality path already covered it,
and the DSL-wide invariant test accepts "compiles OR raises FilterError",
so a 422 passed. _coerce_operand's docstring already stated the contract
its caller wasn't honoring — "Callers handle None (a null check) and `*`
(a wildcard) before calling" — so the guard restores that rather than
adding a new rule.
The three null forms now agree: {"col": null} is IS NULL, {"col": {"ne":
null}} is IS NOT NULL, NOT [{"col": null}] is (IS NULL) IS NOT true.
Also from review of #947:
- Log filter keys, not the body. That log line is new in this branch and
operands carry peer/session ids and free-text `contains` values; the
traceback plus the entry shape is what locates a builder bug.
- Assert whereclause in the _where test helper, so a dropped condition
fails instead of returning the whole statement to substring-match.
- Cover the raw-key JSONB path via source_ids, a JSONB column outside
JSONB_COLUMNS reachable through Document's raw-key fallback.
- Drop the orphaned comment left above ENUM_COLUMN_VALUES when
_coerce_operand replaced SCALAR_OPERAND_TYPES.
- Document that a JSONB column takes an object bare or under `contains`
and nothing else. Bare {"metadata": X} is containment, so the `ne` this
branch removed was never its inverse: a row with {"status":"done","x":1}
satisfied both it and {"metadata": {"status":"done"}}. Per-key operators
and NOT cover the real intents.
- Rewrite "Negation and Unset Fields" to lead with the operator rule and a
truth table, forward-linking to Filtering Conclusions instead of using
conclusions ~525 lines before they are introduced. A conclusion's
session_id is the only nullable documented filterable field, verified
across Message/Document/Session/Peer/Workspace.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: reserve scope__ peer namespace with kind flag and guardrails
Introduce the scope peer namespace (scope__<name>) and the authoritative
{"kind": "scope"} configuration flag, plus the server-side guardrails:
- src/utils/scopes.py: single source of truth for the prefix, kind flag,
and name helpers (scope_peer_name / is_scope_peer_name /
scope_name_from_peer / validate_no_scope_peer_names)
- reject reserved-prefix names on peer get-or-create (422)
- reject scope peers as message authors in crud.create_messages (422)
- reject scope peers as chat/representation targets (422); a scope peer
as the path-level observer is deferred to Phase 2b
- reject scope peers on the generic session-peer add/set/remove routes
and the session-create peers mapping (422, directing to scopes routes)
- peers.list excludes scope peers by default; new PeerGet.kind option
("scope" | "all") switches the view via a configuration JSONB filter
- schemas: Scope / ScopeCreate / ScopeSessions(Add) and
SessionCreate.scopes (unprefixed scope names, validated)
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add scopes CRUD routes and session-create scopes wiring
New /v3/workspaces/{workspace_id}/scopes facade (workspace-level auth;
peer- and session-scoped keys are rejected):
POST "" create-or-get (201/200)
POST /list paginated scope list
GET /{scope_id} single scope
POST /{scope_id}/sessions add memberships
DELETE /{scope_id}/sessions/{session_id} remove membership
GET /{scope_id}/sessions list member session ids
- crud/scope.py: get_or_create_scopes stamps the backing peer with
{"kind": "scope", "observe_me": false} and refuses to adopt a
legacy peer occupying the reserved name without the flag (409)
- memberships are session_peers rows with observe_others=true /
observe_me=false — identical to a hand-built observer peer
- SessionCreate.scopes: create-or-get each scope peer and add the
membership at session creation (the no-backfill common path)
- crud/session.py: public upsert_session_peers wrapper so the facade
bypasses the route-level guardrails without reaching into privates
Backfill of pre-existing documents and reconciliation on removal land in
DEV-1999; membership only affects messages ingested after the change.
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: cover scopes facade, guardrails, and observer semantics
- create-or-get idempotency, list/get, name validation, legacy-collision
rejection (409), auth scoping (workspace key ok, peer/session keys 401)
- reserved prefix rejected on peer create; peers.list kind filtering
- scope peers rejected as message authors, chat/representation targets,
and on the generic session-peer routes
- membership add/list/remove with observe_others=true / observe_me=false
row shape asserted via DB, and facade-less equivalence with a
hand-built observer peer
- end-to-end litmus: after adding a session to a scope, the deriver
enqueue fan-out includes the scope peer as an observer
- session creation with scopes: [a, b] creates both memberships
Part of DEV-1997 (Scopes RFC DEV-1970).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(crud): preserve cache invalidation across get_or_create retry
`get_or_create_peers` and `get_or_create_scopes` mutate existing rows, then
insert new ones inside `db.begin_nested()`. When a concurrent writer creates
one of those rows first, the insert raises IntegrityError and the function
retries.
`begin_nested()` autoflushes the pending UPDATEs *before* opening the
savepoint, so the rollback neither undoes them nor expires the now-clean ORM
state. The retry then compared already-updated values, found no change, and
dropped those peers from `changed_peers` — skipping the cache purge while the
row change committed anyway, leaving entries stale until the 300s TTL.
Carry the mutated names into the retry via `_pending_invalidation` so the
purge cannot be lost.
The scopes facade mirrors `get_or_create_peers`, so both copies carried this.
The peer path is pre-existing and runs on every message ingest.
Also add the missing /v3 prefix to `_SCOPES_ROUTE_GUIDANCE`, which pointed
callers at a 404.
Adds tests/crud/test_get_or_create_retry_invalidation.py, which drives a real
racing session and fails without this change.
* fix(scopes): make scope identity unforgeable, unblock non-pattern peer names
Three coupled changes to the scopes facade.
1. `PeerCreate` no longer gates internal lookups. It exists to validate a new,
user-supplied peer id at the API boundary, but crud used it as a DTO for
names that already exist, so any name outside RESOURCE_NAME_PATTERN raised a
raw pydantic ValidationError — which is not a HonchoException, so it fell
through to the catch-all handler as an HTTP 500. Adds `PeerSpec` (same
fields, no charset pattern) as `PeerCreate`'s base, widens
`get_or_create_peers` to accept it, and changes `get_peer` to take a plain
str. All 13 construction sites converted; the create route keeps full
validation.
This unbreaks the Dreamer: DreamScheduler passes `collection.observer`
straight into the specialist preflight, and scope peers have
`observe_others=true`, so every `(scope.x, peer)` dream died there — the
feature scopes exist to enable. It also fixes a pre-existing bug unrelated
to scopes: a peer named `alice.smith` (legal before d429de0e5338, which
validated names by length alone) 500s on message create, session peer add,
and peer update.
2. The `kind` flag moves from `configuration` to `internal_metadata`.
`configuration` is user-writable — `PeerCreate`/`PeerUpdate` accept a
free-form dict and `update_peer` replaces it wholesale — so a legitimate
`{"observe_me": true}` update silently dropped the flag, and a forged
`{"kind": "scope"}` injected an ordinary peer into `POST /scopes/list`.
`internal_metadata` appears in no API schema. `observe_me: false` stays in
`configuration`, where it belongs.
3. Scope identity requires prefix AND flag, via `is_scope_peer()` and
`scope_peer_clause()`. Neither half is forgeable: the prefix sits outside
RESOURCE_NAME_PATTERN, `internal_metadata` is unreachable. Usage-site guards
become flag-based so a legacy peer merely occupying the namespace keeps
working rather than 422-ing on its own traffic; peer create and update stay
name-based, since those must stop new names entering the namespace.
`update_peer` now returns 422 instead of 500.
Also swaps the reserved prefix from `scope__` to `scope.`: `_` is inside
RESOURCE_NAME_PATTERN, so any tenant could already own a `scope__x` peer.
No DB migration — `internal_metadata` already exists on `peers`, and no scope
peers exist in any deployment yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): validate peer names on create, close namespace squatting and upsert race
Addresses three review findings against 48047a6a.
1. `PeerSpec` let API callers create invalid and reserved-prefix peers.
Widening `get_or_create_peers` to accept a pattern-free schema fixed the
lookup 500s but also removed validation from the *insert* path, and
request-controlled names reach it via message authors, session peer maps, and
the chat observer path — none of which carry a charset pattern of their own.
Confirmed: `POST /sessions/{id}/messages` with `peer_id: "scope.x"` returned
201 and minted an unflagged squatter, after which `POST /scopes {id: x}` was
permanently 409-blocked — namespace denial of service by any caller able to
post a message. `peer_id: "not a valid name!@#"` was likewise created.
Fixed by validating only names about to be INSERTed
(`_validate_new_peer_names`), so already-existing names — legacy dotted
names, scope peers — still resolve without a spurious 422. That keeps the
Dreamer fix intact, since it reads through `get_peer`.
2. Existing reserved-prefix squatters could not be updated. The name-based guard
on `PUT /peers/{peer_id}` refused every `scope.` name, contradicting the
invariant that an unflagged squatter stays a normal peer. Now flag-based, so
behavior is three-way: a real scope is refused, an existing unflagged peer
updates, and a missing reserved-prefix name is refused by (1) rather than
minted.
3. Scope checks raced with get-or-create and the membership upsert. The
route-level guards run before peers are resolved, so a scope created
concurrently in that window would be attached by the generic path with a
default `SessionPeerConfig()`, clobbering its observer membership config.
Adds `_reject_resolved_scope_peers`, which runs on the resolved rows in the
same transaction as the upsert — no window, no extra query. The early checks
stay for better error messages.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): guard scope membership config, move checks to the mutation point
Addresses a second review pass against 10655792.
1. Scope membership configuration was directly user-mutable.
`PUT /sessions/{id}/peers/{peer_id}/config` had no scope guard at all, and
`crud.set_peer_config` resolved the peer only to discard the row. Confirmed:
posting `{"observe_others": false, "observe_me": true}` for a scope returned
204 and persisted, which silently stops all fan-out into the scope and makes
Honcho form a representation *of* a scope — neither of which is reachable
through the facade. Deterministic, no race required. Now checked on the row
`get_peer` already returns, so it costs nothing and cannot race.
2. Empty and over-long names were still 500s. Removing the charset pattern from
`PeerSpec` fixed one trap but left its length bounds, and request-bound peer
names carry no length limits of their own — so `peer_id: ""` or a 513-char
name reached `PeerSpec(...)` and raised a raw pydantic ValidationError that
the catch-all turned into a 500. `PeerSpec` now carries no constraints at all
(matching its documented purpose) and every rule for a new name lives in
`_validate_new_peer_names` on the insert path.
3. Resolved-row protection generalized. The previous pass applied it only to
membership upserts, leaving check-then-use windows elsewhere: peer update
could have a concurrently-created scope's configuration replaced wholesale
(create-path validation does not fire for a peer that now exists), the chat
observer get-or-create could resolve a fresh scope as its observer, and the
generic session-peer removal could silently detach a scope from its sessions.
Each now inspects the resolved peer immediately before acting; the redundant
name-level guard on the update route is dropped in favor of the race-free one.
`remove_peers_from_session` grows an internal `_allow_scope_peers` flag because
the scopes facade ends membership through that same path and must not be blocked
by its own guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(scopes): enumerate every peer-touching route against a scope policy
Three review passes each found the same class of defect: a route nobody had
checked, rather than logic that was subtly wrong. One of them — `PUT
/sessions/{id}/peers/{id}/config`, which let any caller set a scope to
`observe_others=false` and silently stop all fan-out into it — predated this work
entirely, because the guardrail set was assembled guardrail-by-guardrail instead
of derived from the route list. Sampling review cannot close that kind of gap;
enumeration can.
Derives every route through which a peer name can reach the system and requires
each to be classified as GUARDED or EXEMPT-with-a-reason, so a newly added
peer-touching route fails the suite until someone classifies it. Detection is the
union of path shape and a walk of the dependant tree (including sub-dependency
`Form(...)` params and nested request-body models), because neither signal alone
suffices: parameter names miss `POST /sessions/{id}/peers`, whose peer names are
dict keys, and path shape misses `messages/upload`, whose `peer_id` arrives as a
form field behind a parser dependency.
Both invariants are then asserted behaviorally, by calling the routes rather than
inspecting annotations — the guards deliberately live in crud, which is what makes
`messages/upload` guarded for free via `crud.create_messages`:
- a real scope is refused on all 11 guarded routes, and the rejection must name
the scope, so an unrelated 422 (a malformed body) cannot pass the assertion;
- an *unflagged* peer merely occupying the reserved namespace is unaffected. That
half regressed once already when `update_peer` used a name-based check.
Mutation-tested all three failure modes: disabling the `set_peer_config` guard
fails the guarded test naming that route; regressing `update_peer` to name-based
fails the squatter test; adding an unclassified peer route fails the enumeration.
Covers the HTTP surface only. Peer names also reach the system through the
deriver, dreamer, and queue, which have no route table to enumerate — noted in
the module docstring rather than implied to be covered.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): refuse a scope in every observed position; enumerate per position
Addresses a fourth review pass against 62681e4d. The headline finding is that the
previous commit's enumeration test had the wrong *model*, not a missing entry.
1. Manual conclusions could create knowledge about a scope. `POST /conclusions`
validated only that observer_id and observed_id exist, so a scope as
`observed_id` persisted a conclusion about a peer carrying observe_me=false and
created an (observer, scope) collection for it. Confirmed: 201, and it read
back. `POST /schedule_dream` had the same hole via `observed`.
The fix is positional, because the invariant is:
A scope may be an OBSERVER. A scope may never be OBSERVED.
A scope as `observer_id` is how scoped conclusions are stored and must keep
working (verified still 201); as `observed_id` it is now refused. Same split
applied to schedule_dream `observed`, the peer-card `target` (which also covers
a scope's self-card, since target-omitted collapses observed to peer_id), and
session-context `peer_target`.
2. Chat target and both representation roles kept check-to-use races. Only the
chat path-level observer was re-checked on its resolved row; the target was
checked by name and then resolved without inspecting scope identity. Both are
now checked at the dialectic preflight, where observer and observed are already
resolved — an absent name has already failed by then, and an existing squatter
cannot retroactively become a scope.
3. Generic membership removal was still racy. The adjacent SELECT narrowed the
window but could not close it under READ COMMITTED. The UPDATE now carries its
own correlated NOT EXISTS against scope_peer_clause(), so Postgres evaluates
the exclusion as part of the statement and a scope committed after the advisory
check still cannot be detached.
4. New-name validation ran after the name reached Postgres. A NUL byte passed the
request schemas and PeerSpec, then raised psycopg.DataError inside the lookup —
a 500. Values that cannot correspond to a stored row by construction (NUL
bytes, over-length names) are now refused before the query.
(Over-length names already returned 422; only the wasted query was real there.)
The enumeration test is rekeyed from (method, path) to (method, path, position).
A binary per-route verdict cannot express finding 1 at all: `POST /conclusions` is
one route with two positions and opposite verdicts. Detection widens to observer /
observed / target / peer_target / peer_perspective, which surfaced four routes the
previous version never saw — conclusions, schedule_dream, queue/status, and
session context.
Also registers `src.routers.workspaces.tracked_db` in the conftest patch list; the
new guard there would otherwise have run against the real configured database
instead of the per-test one.
Mutation-tested: disabling the conclusions observed-guard fails the positional
test naming that position; adding an unclassified `observed_id` param fails
enumeration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): refuse future scopes in observed positions, preserve scope membership
Addresses a fifth review pass against 14136e5b. All six findings reproduced
locally before fixing.
1. High — generic peer replacement removed scope memberships.
`set_peers_for_session` soft-deleted every active SessionPeer row, and the
request-level guard only inspected names *present* in the replacement map. A
caller detached a scope by simply omitting it, never naming it — so no
request-level guard could ever see it. Reproduced: scope sessions went
['<id>'] -> [] on a 200. The exclusion now lives in the UPDATE itself
(correlated NOT EXISTS against scope_peer_clause), so replacement means
"replace ordinary peers" regardless of request contents or concurrent creation.
2. High — peer cards could be pre-seeded for future scopes.
`set_peer_card` resolves only the observer and writes a JSONB key derived from
an unchecked observed name, and the route guard rejected only *existing*
flagged scopes. Reproduced: PUT card with target=scope.<missing> returned 200,
creating that scope then returned 201, and the card described the real scope.
3. High — dreams could be queued for future scopes.
The route checked `observed` in a read-only session that closed before
`enqueue_dream`, and a missing reserved name passes any is-it-a-scope check.
Reproduced: 204 with observed=scope.<missing>.
2 and 3 share a root cause, so they share a fix: a new `reject_scope_observed`
that is stricter than `reject_scope_peers` in exactly one case — a *missing*
reserved name is refused, because nothing on these paths creates the peer, so
nothing else would ever catch it. Existing unflagged squatters still pass.
Both guards moved to the mutation point: card validation into
`crud.set_peer_card` (same transaction as the JSONB write, so Dreamer and
agent-tool callers are covered), dream validation into `enqueue_dream` (same
transaction as the queue insert). The redundant route-level checks are dropped
rather than left as weaker duplicates.
4. Medium — prefixed NUL names still reached PostgreSQL.
`reject_scope_peers` filtered for the reserved prefix and sent matches to a
text comparison, so "scope.future\0name" raised psycopg.DataError — a 500.
Both guards now share `_reserved_name_candidates`, which materializes the input
once and rejects impossible values before any SQL. Materializing matters
independently: the message-author path passes a generator, and validation
iterates separately from the prefix filter, so a generator would be
half-consumed. `_reject_impossible_peer_names` now takes a Collection so the
type checker enforces that.
5. Medium — representation kept a check-to-use race.
The previous commit claimed both representation roles were rechecked after
resolution; that was wrong — only the dialectic preflight got that check, and
the representation route never goes through it. It now opens one short
read-only session *after* the embedding call, checks both positions, and passes
that same session to `get_working_representation`, so no connection is held
across external work and a scope committed later cannot have conclusions in the
collection being read.
6. Low — policy coverage was not exhaustive. `sender_id` reaches CRUD as
`observed` but was missing from the detected parameter set. ALLOW cases could
also not carry builders, so the suite never proved the other half of the
contract — that legitimate scope *observers* keep working, which a guard
rejecting scopes everywhere would satisfy. Both fixed; observer positions on
conclusions, dreams, cards, session context and queue status are now asserted
behaviorally.
Deliberately not implemented: the scope-creation backstop scanning for
pre-existing card keys and queue items naming a future backing peer. Reasoning is
recorded in `get_or_create_scopes` — no new such state can be created now, any
pre-existing row is coincidental since `scope.` was never a meaningful namespace,
the consequence is inert, and detecting card keys means a full table scan per
scope creation.
Mutation-tested each new guard: removing the replacement exclusion fails both
membership-preservation tests; weakening either observed guard to existing-only
fails the pre-seeding tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): exclude scope memberships from the session observer limit
Scope memberships carry observe_others=true, so every scope counted against
SESSION_OBSERVERS_LIMIT (default 10) — capping scopes-per-session at the limit
minus the session's real observers, and reporting the failure as
`400 Cannot create session <name> with 11 observers. ... Observers are peers
with 'observe_others' set to true.` on a membership call. Wrong on three counts:
the ceiling is undocumented and contradicts RFC §5.1 ("sessions belong to any
number of scopes"), the message describes session creation, and it leaks the
word "observer" through a facade whose entire job is hiding observers (RFC
goal 5). The limit exists to bound per-observer deriver fan-out for real peers;
a scope costs document rows, not LLM calls (RFC §5.2), so it does not belong in
that budget.
Excluded from both halves of the check in `_get_or_add_peers_to_session`: the
incoming names via a flag-based lookup, existing memberships via a correlated
NOT EXISTS on `scope_peer_clause()` — the same pattern the replacement and
removal paths already use, so the exclusion holds regardless of concurrent
scope creation. The early `count_observers_in_config(session.peer_names)` check
in `get_or_create_session` is left alone: `peer_names` cannot contain a scope,
and `scopes` is a separate field.
`reject_scope_peers` is split into a `scope_peer_names()` query helper plus a
two-line raiser so the observer count reuses the authoritative name-AND-flag
predicate instead of growing a third copy of it. Still costs nothing on the
common path — no reserved-prefix name in the input means no query at all.
Also caps `SessionCreate.scopes` at 100, matching `ScopeSessionsAdd.session_ids`.
This belongs in the same commit: the observer limit was the only thing bounding
that list, so removing it turns an unbounded `scopes` array into a peer row and
a membership row per element, committed — the single-request path to the
cardinality anti-pattern RFC §8 warns about. Partly answers OQ6: no per-session
cap, 100 per request.
Tests: a session joins SESSION_OBSERVERS_LIMIT + 2 scopes through both the
facade and session creation; real observers over the limit still 400, so the
carve-out cannot quietly disable the limit; 101 scopes is a 422.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): skip semantic retrieval when the embedding precompute failed
Addresses three open review comments.
1. Major — the representation read could embed inside its DB session. The
route's precompute is suppressed, and both
`RepresentationManager.get_working_representation` and
`crud.query_documents` fall back to embedding when a query arrives without
one, so a failed precompute meant an external call inside the read session
this branch opens for the scope re-check — the connection-holding rule the
route's own comment claimed to satisfy. The innermost fallback also only
catches ValueError, so a provider outage surfaced as a 500. The semantic
query is now passed only when an embedding exists, degrading to
derived+recent retrieval. (`crud.query_documents` embedding inside a caller's
session predates this branch and is left alone.)
2. Minor — `test_resolved_scope_peer_rejected_at_membership_upsert` described a
race it does not perform. It creates an already-flagged scope and calls crud
directly; the unflagged → flagged transition is not simulated. Docstring now
says what the test actually pins.
3. Minor — `test_empty_replacement_preserves_scope_membership` asserted only
half its docstring. It passed if the empty PUT left every ordinary
membership intact; now asserts the ordinary peer's left_at is set.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scopes): close auth and observed-position gaps, paginate membership
Review response for #884.
Security:
- gate `SessionCreate.scopes` behind a workspace-level key; the session-create
route is self-authorizing, so a peer- or session-scoped token could mint scope
peers and join sessions to scopes it had no access to via `POST /scopes`
- refuse a reserved-but-nonexistent name in two observed positions that used the
permissive guard: chat `target` and session-context `peer_target`. Both let a
caller act on `scope.X` before it existed, then create the scope
Facade:
- exclude scope peers from `GET /sessions/{id}/peers` and refuse the membership
-config read for a real scope, matching its write side
- replace `GET /scopes/{id}/sessions` with `POST /scopes/{id}/sessions/list`
returning `Page[Session]`; the add route now returns 204. Membership was
unbounded on both, while every other list surface paginates
- rename `crud.get_scope` to `get_scope_or_raise`
Tests:
- add a missing-name axis to the route-policy table (`Case.refuse_missing`), which
is what surfaced the two guard gaps above
- delete 14 hand-written tests the table now enumerates; 52 -> 39 functions in
test_scopes.py with more cases covered
- tighten the squatter assertion from `!= 422` to `< 400`, which was passing on 5xx
- assert the FastAPI-internals traversal still derives positions, so a framework
upgrade can't silently empty the suite
Docs:
- drop internal ticket and RFC references from the published OpenAPI descriptions
and surrounding comments; state the behavior instead
- move implementation reasoning out of the `PUT /peers/{id}` docstring, which
FastAPI publishes, into a comment
* test(scopes): assert exact statuses for permissive missing-name cases
Follow-up review pass on #884.
- add `Case.missing_status` so a permissive missing-name position asserts the
status it should actually get (404, or 200 for the no-op removal) instead of
`!= 422`, which also passed on a 5xx — the same hole already closed in the
squatter assertion
- require it whenever `refuse_missing` is False, and require its absence when
True, so the policy table can't drift from the assertion
- repoint a stale allow-reason at POST /scopes/{scope_id}/sessions/list; the GET
it named was removed
- document the membership list's ordering under `reverse`
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix: apply session scoping to all working-representation query paths
session_name was only applied to the recent-documents query in
RepresentationManager; the semantic and most-derived paths ignored it,
so limit_to_session leaked cross-session conclusions into perspectives.
- Thread a session allowlist (session_names) uniformly through all
three query paths; pushed down to pgvector and external vector stores
- Accept a list so the upcoming session-allowlist API reuses this path
- Fail closed on an empty allowlist (downstream stores drop empty IN
clauses, which would silently widen scope)
Fixes DEV-1994
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: bare-list membership sugar in the filter DSL
{"session_id": ["s1", "s2"]} is now shorthand for
{"session_id": {"in": [...]}} on regular columns, generically
(peer_id, etc.). JSONB metadata columns are excluded — a bare list
there keeps JSONB containment semantics, unchanged.
Previously a bare list on a regular column compiled to a type-mismatched
equality that matched nothing, so this is strictly additive.
Also translates the same shape in the turbopuffer/lancedb filter
builders, and fixes lancedb dropping empty IN clauses (fail-open) —
an empty membership list now emits an always-false condition.
Groundwork for DEV-1995 (session allowlist via the existing filters
DSL, no new API params)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: session allowlist on dialectic and representation via filters
Adds a constrained 'filters' body to peer.chat and /representation —
the same DSL search and conclusions already accept, supporting only the
session_id key (a session id, a bare list, or {"in": [...]}).
Unsupported keys and shapes are rejected with 422, never silently
ignored. Composes with session_id (must be included in the allowlist
when both are given). Capped at 1,000 sessions per request.
Enforcement is uniform at every recall chokepoint, fail-closed:
- dialectic prefetch + search_memory: conclusion recall restricted to
the allowlist; dream docs (session_name IS NULL) excluded
- message tools (search/grep/date-range/temporal/context/history):
strict intersection of allowlist and observer session membership
- get_reasoning_chain: unavailable under an allowlist (chains traverse
provenance across sessions and cannot be scoped without leaking)
- empty allowlist short-circuits to empty results everywhere
Auth: workspace keys pass the allowlist as-given; peer-scoped JWTs must
be a member of every allowlisted session (403 otherwise), mirroring the
existing single-session check.
Fixes DEV-1995
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: fail closed on empty session allowlist across all filter builders
Empty session allowlists relied solely on the early-return guard in
_get_working_representation_internal. The layers below it were
inconsistent, so a future direct caller (the DEV-1995 allowlist API)
would silently widen scope instead of failing closed:
- _build_filter_conditions used a truthiness check; an empty list was
treated like None and dropped the filter. Now uses `is not None`,
matching the recent/most-derived SQL paths.
- turbopuffer emitted a bare `In []` with undocumented (possibly
fail-open) semantics. Now emits an explicit always-false predicate,
mirroring lancedb's `1 = 0`.
Also extract the duplicated JSONB column tuple in filter.py to a
JSONB_COLUMNS constant.
Tests exercise each fail-closed guarantee at the layer it lives, rather
than masking it behind the early-return guard.
* fix: address tests
* fix(crud): fail closed when session_name is outside the allowlist
search/grep/history helpers scoped to a single session_name ignored the
session_names allowlist entirely — a caller could read a session the
allowlist forbids. The API routes guarded this with a 422, but the
dialectic tools call these CRUD functions directly and bypassed it.
Enforce it at the boundary: return [] when session_name is set and not in
the allowlist, across _semantic_search_messages (covers search_messages +
search_messages_temporal), grep_messages, get_messages_by_date_range,
get_recent_history, and get_observation_context.
Also rename the public param allowed_sessions -> session_names for
consistency with representation.py / chat.py / peers.py; the resolved
intersection keeps its distinct name allowed_session_names.
* fix: test
* fix(scopes): tighten and consolidate session allowlist per review
Addresses review feedback on the session allowlist (DEV-1995).
Behavior changes:
- Auth gate on peers.chat now uses active membership (left_at IS NULL)
via get_peer_session_names(active_only=True), matching the adjacent
is_peer_in_session check on options.session_id. Previously a peer that
had left a session was denied when naming it directly but permitted
when naming it in filters.session_id.
- Scoped conclusion recall is restricted to level == "explicit"
(ALLOWLIST_SAFE_LEVELS). Dream-derived conclusions are stamped with a
single session_name but synthesized across all sessions, so that stamp
can't be scoped on. Applied at all four recall paths. Unscoped recall
is unchanged. Follow-up to give conclusions an authoritative
source-session set is tracked in DEV-2201.
- The allowlist gate checks `is not None` rather than truthiness, so
filters={"session_id": []} reaches it instead of being skipped.
Refactors:
- New crud.message.resolve_session_scope replaces four near-identical
copies of the allowlist-membership intersection. Returns
(allowlist, deny) and never returns an empty list, so the None vs []
distinction that external stores fail open on lives in one tested
place. Takes db=None and opens its own short-lived session only when
distinction that external stores fail open on lives in one tested
place. Takes db=None and opens its own short-lived session only when
an observer lookup is needed, preserving external-lookup-first
ordering on the vector-store path.
- extract_session_allowlist takes must_include, collapsing the
session_id-in-allowlist check duplicated across both peer routes.
- DialecticAgent._select_tools dedupes the two toolset-selection blocks
and drops get_reasoning_chain under an allowlist, rather than paying
for the schema plus a wasted turn to return a refusal.
- Rename session_names -> session_allowlist across crud, agent tools,
dialectic and routes, to remove the one-character ambiguity with
session_name. Internal only; the public filters.session_id surface is
unchanged.
Docs:
- session_allowlist documented across all message and recall entry
points, including the None / [] / populated contract.
- session_name marked deprecated for scoping. Not removed and not
aliased: it also pins the query to one session, bypasses observer
scoping, and drives session-history injection into the dialectic
prompt, so it has no drop-in replacement.
- Note at the Document branch in utils/filter.py that the raw-key
fallback is load-bearing for session scoping.
Tests: 20 -> 39 in tests/test_session_allowlist.py, covering the
peer-scoped JWT gate (member, non-member, left-session, workspace-key
bypass, empty allowlist), the resolve_session_scope tri-state including
the no-DB-checkout path, must_include, and the level narrowing.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix: apply session scoping to all working-representation query paths
session_name was only applied to the recent-documents query in
RepresentationManager; the semantic and most-derived paths ignored it,
so limit_to_session leaked cross-session conclusions into perspectives.
- Thread a session allowlist (session_names) uniformly through all
three query paths; pushed down to pgvector and external vector stores
- Accept a list so the upcoming session-allowlist API reuses this path
- Fail closed on an empty allowlist (downstream stores drop empty IN
clauses, which would silently widen scope)
Fixes DEV-1994
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: bare-list membership sugar in the filter DSL
{"session_id": ["s1", "s2"]} is now shorthand for
{"session_id": {"in": [...]}} on regular columns, generically
(peer_id, etc.). JSONB metadata columns are excluded — a bare list
there keeps JSONB containment semantics, unchanged.
Previously a bare list on a regular column compiled to a type-mismatched
equality that matched nothing, so this is strictly additive.
Also translates the same shape in the turbopuffer/lancedb filter
builders, and fixes lancedb dropping empty IN clauses (fail-open) —
an empty membership list now emits an always-false condition.
Groundwork for DEV-1995 (session allowlist via the existing filters
DSL, no new API params)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: fail closed on empty session allowlist across all filter builders
Empty session allowlists relied solely on the early-return guard in
_get_working_representation_internal. The layers below it were
inconsistent, so a future direct caller (the DEV-1995 allowlist API)
would silently widen scope instead of failing closed:
- _build_filter_conditions used a truthiness check; an empty list was
treated like None and dropped the filter. Now uses `is not None`,
matching the recent/most-derived SQL paths.
- turbopuffer emitted a bare `In []` with undocumented (possibly
fail-open) semantics. Now emits an explicit always-false predicate,
mirroring lancedb's `1 = 0`.
Also extract the duplicated JSONB column tuple in filter.py to a
JSONB_COLUMNS constant.
Tests exercise each fail-closed guarantee at the layer it lives, rather
than masking it behind the early-return guard.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix: enforce explicit-document session purity in dedup/merge paths
Audit for DEV-2000 (Scopes RFC prerequisite): explicit-level documents must
stay session-pure so scope memory can be built by copying explicit documents
between collections. Two classes of violation were possible:
- Exact-content and semantic dedup in crud/document.py matched candidates
with no level or session scoping, so an explicit document could be
reinforced by — or soft-deleted in favor of — a same-content document from
a different session or a different level (silently merging cross-session
derivations into one row).
- The generic create_observations tool handler accepted level='explicit'
from agents with no message context (dreamer/dialectic), which would mint
session-less explicit documents.
Enforcement (refuse, never rewrite):
- create_documents refuses explicit documents with a null session_name
- exact dedup keys on (content, level, session-for-explicit); derived levels
keep cross-session consolidation
- is_rejected_duplicate scopes candidate search to the same level, and the
same session for explicit documents
- the create_observations tool rejects explicit-level input outside message
ingestion (deriver) context
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: add card_refresh dream type for event-driven peer-card updates
Adds a lightweight dream variant (DEV-2000, Scopes RFC prerequisite) that
runs ONLY the peer-card update — for event-driven refreshes such as scope
membership changes and cold starts:
- DreamType.CARD_REFRESH alongside OMNI; dispatched by process_dream to a
new run_card_refresh_dream orchestration
- CardRefreshSpecialist: restricted to get_recent_observations,
search_memory, and update_peer_card (no observation-mutating tools), with
a low tool-iteration cap of min(6, DREAM.MAX_TOOL_ITERATIONS)
- rebuild=True mode carried in the dream payload: the existing card is NOT
injected into the prompt and the specialist rebuilds it solely from
observations present in the collection (for use after removals)
- enqueue-able via the manual enqueue_dream path (bypasses volume gates);
the work-unit key already embeds the dream type so a card refresh never
collides with a pending omni dream. POST /v3/workspaces/{id}/schedule_dream
accepts dream_type=card_refresh plus the rebuild flag
- card refreshes never advance the omni dream guard pair
(last_dream_at / last_dream_document_count)
- shared PEER CARD prompt section extracted (verbatim) from
DeductionSpecialist for reuse; CallPurpose gains dream.card_refresh
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: fix tests
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Structured outputs for dialectic
* cleanup
* rename json_schema_to_pydantic to clarify it's not a general schema converter
* clean up schema DoS guards
* simplification and cleanup of schema conversion
* chore: ruff and pyproject toml
* chore: basedpyright cleanup in test
* fix: some needed unrelated test failures
* test(schema_conversion-and-anthropic-backend): expand test coverage
include table tests
* fix(llm): support combined tool calling and structured output across backends
- OpenAI: parse() 500s on non-strict function tools; route tool-carrying
structured requests through create() with an explicit json_schema
response_format (mirrors the streaming path)
- Anthropic: skip the '{' JSON prefill when tools are present so tool_use
blocks stay reachable; make the schema instruction conditional and rely
on parse + repair
- Gemini: native response_schema + function calling is rejected before
Gemini 3; with tools present, inject a schema instruction into the final
turn instead and rely on parse + repair
- All backends: tool-call turns carry no consumable content, so skip
structured-output parsing on them
Extracted from the dialectic structured-output branch (DEV-1652) so the
transport layer can land independently.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(live_llm): exercise combined tools + structured output per provider
Two-turn live flow per backend: a forced tool-call turn (structured
parsing must be skipped) followed by a replay turn that must return a
schema-conforming answer with tools still attached. Asserts the
provider-specific request shaping: no parse() for OpenAI (500s on
non-strict tools), no '{' prefill for Anthropic, no native
response_schema for Gemini.
Verified against live OpenAI (gpt-4.1, gpt-5, gpt-5.4, gpt-5.4-mini)
and Gemini (gemini-2.5-flash).
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(unified): dialectic chat with response_format schema under tool use
Adds response_format pass-through to the unified runner's chat query and
a test case that forces the dialectic tool loop (reasoning off + global
enumeration question) while requiring a schema-conforming JSON answer —
end-to-end coverage of the combined tools + structured output transport
path on whichever provider each level is configured with.
Verified locally against a full harness run (json_match assertions pass;
the llm_judge assertion additionally runs in CI where the Anthropic key
is available).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: some needed unrelated test failures
* ci: add label-triggered live LLM test workflow
Adding the run-live-llm label to a PR (or workflow_dispatch) runs
tests/live_llm/ against real provider APIs — the only place the
--live-llm suite runs in CI. Reuses the unified-tests environment and
its Secrets Manager staging-dotenv resolution for provider keys; runs
on ubuntu-latest (no Fly runner, no Docker — the suite only touches the
LLM backends). Pins LIVE_LLM_ANTHROPIC_45_PLUS_MODELS=claude-sonnet-4-5
since the Anthropic family has no default models and would otherwise
silently collect empty.
Opt-in by design: live model behavior is variable, so this is a signal,
not a required check.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: run live LLM tests on main pushes touching the transport
Mirrors unified-tests' push trigger, scoped to paths that can affect
the live suite (src/llm/, config, the tests, deps, and the workflow
itself) so provider API calls aren't spent on unrelated changes.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: disable auth in live LLM test environment
The staging dotenv sets AUTH_USE_AUTH=true without a usable JWT secret,
and src/config.py validates the pair at import time — the same reason
unified-tests overrides it. This suite never runs the API server.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(live_llm): fix gpt-5.4 reasoning_effort and gemini replay-turn flake
- test_live_openai: gpt-5.4 dropped 'minimal' from the reasoning_effort
vocabulary, so the gpt5 caching test 400'd — and the OpenAI backend's
BadRequestError terminal swallowed it into an empty CompletionResult.
Pick the effort per model generation.
- test_live_tools_structured_output: use tool_choice='auto' on the
replay turn, matching the production dialectic loop (which never
forces 'none') — NONE mode is what provoked gemini-2.5-flash's empty
candidates. Drop the temperature pin so retries actually resample,
and treat a repeat tool call as a retryable attempt.
Verified live: full suite green, gemini 4/4 consecutive passes.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: fail live LLM run when no staging secret was loaded
If the latest-tag fetch fails and no second tag exists, the fallback
step is skipped rather than failed, and the job would proceed without
provider keys — every test then skips via require_provider_key and the
run goes green. Guard on both fetch outcomes so that path fails loudly.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(live-llm-tests-GHA): remove extra comments
* feat(structured-output): enable non-recursive schema references
* docs(structured-outputs): clean up new doc
* test(structured-output): fix caching refs memory leak, add tests
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(telemetry): CloudEvents + Langfuse tracing as projections over a captured LLM stream
Capture each LLM call once (CapturedLLMCall) and fan it out to multiple
exporters -- "one data model, two projections": a CloudEvents trace stream
(llm.call.traced / trace.content) and a Langfuse projection, both reconstructing
trace -> run -> step -> generation from the same source of truth.
- Capture seam (src/llm/capture.py): one canonicalization + content-addressed
hashing point, with an O(N) per-span memo so repeated context isn't re-hashed.
- Session correlation threaded telemetry -> captured call -> exporters,
namespaced only at the Langfuse export boundary.
- Span identity consolidated onto LLMTelemetryContext; dropped TRACE_ENDPOINT.
- Canonical generation/step names; dreamer branches nest under one dream trace;
tool calls become spans under their step.
- LANGFUSE_EXPORTER_MODE toggle ("exporter" default; "inline" kept one release
for side-by-side validation), centralized into computed settings predicates.
- Per-run/per-trace dedup registries (trace_session, langfuse_session) bounded
by an LRU so dedup and span grouping survive long-running workers.
- Embedding-call tracing; deterministic high-volume event sampling.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(telemetry): address trace-review findings (span/step_seq collisions, test, logging)
- Dreamer specialists mint a distinct span_id per execution (trace_id stays the
shared dream run_id), so their CloudEvents trace resource ids no longer collide
between deduction and induction.
- Tool-loop no-tool early-return streams the tail with the next ordinal
(iteration+2) instead of reusing the in-loop call's step_seq, avoiding a
colliding trace resource id; mirrors the synthesis path.
- Tighten test_clips_oversized_string to assert output stays within TRACE_MAX_BYTES.
- emit_trace logs the swallowed exception with exc_info for debuggability.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(telemetry): silence exporter-mode Langfuse warning + drop summarizer run_id placeholder
Two CloudEvents/Langfuse correctness fixes, independent of the trace viewer.
Langfuse exporter-mode gating: annotate_current_generation_io (and its two
executor.py call-site guards) were gated on LANGFUSE_PUBLIC_KEY instead of
langfuse_inline_enabled. In the default `exporter` mode they called
get_client().update_current_generation() with no active @observe span, logging
"No active span in current context" (~14 per dialectic run) and building
throwaway model_dump payloads on every LLM call. The LangfuseExporter projects
I/O from the captured stream, so these helpers must no-op in exporter mode.
Gated all three on langfuse_inline_enabled; added a regression test; fixed a
stale conditional_observe docstring.
Summarizer run_id placeholder: AgentToolSummaryCreatedEvent hardcoded
run_id="deriver"/iteration=0 because summarization is a single LLM call, not an
agentic run. That placeholder pollutes run_id grouping in the CloudEvents stream
(any consumer that groups by run_id sees a phantom "deriver" run). Made
run_id/iteration optional (None) and re-keyed get_resource_id on
message_id:summary_type (the real per-summary identity; run_id/iteration can no
longer identify it); bumped schema_version 2->3. Xatu ingestion stores only the
CloudEvent envelope, so the field/resource_id/version changes are transparent to it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: update docstrings to be less verbose
* fix(telemetry): address PR review on captured-stream tracing
- embedding traces get a fresh span_id under parent_span_id=run_id, so
sibling embeddings in one run no longer share a span/idempotency key
- capture the provider finish_reason from stream chunks instead of
hardcoding "stop" on a successful drain
- gate the Langfuse exporter behind TELEMETRY.ENABLED (master switch) so
disabling telemetry sends no traces at all
- rename _emit_derived_content -> _emit_hashed_content
- inline the _emit_trace wrapper; drop unused trace_session.end_run
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor: rename TELEMETRY_TRACE_PAYLOADS to TELEMETRY_TRACE_PAYLOADS_ENABLED
* fix(telemetry): capture provider tool calls in trace stream
The captured trace stream dropped assistant tool calls for openai/gemini:
build_captured_messages only read {role, content, tool_call_id}, but those
providers keep tool calls outside content (openai's tool_calls, gemini's
parts), so replayed tool-call turns landed as empty content and gemini lost
its text and tool results entirely. Anthropic (tool_use in content) was fine.
Normalize each input message per provider into a unified tool_calls
[{id, name, input}] field on CapturedMessage/TraceContentEvent, recovering
gemini text/results along the way, and fold tool_calls into
compute_content_hash so empty-content openai turns no longer collide in the
dedup store. langfuse_exporter._input now surfaces the calls.
Also fix a silent serialization drop: gemini thought_signature is bytes, so
model_dump(mode="json") on the traced event raised UnicodeDecodeError and
emit_trace swallowed it -- dropping the whole tool-calling iteration from the
trace stream (billing and Langfuse were unaffected). base64-encode the
signature on the telemetry path; replay keeps the raw bytes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(telemetry): type replay tool-call dict for bytes signature
thought_signature widened to str | bytes | None, but
_tool_call_result_to_dict's literal was inferred as
dict[str, str | dict[str, Any]], so the bytes assignment failed project-wide
basedpyright (the per-file pre-commit hook didn't catch it). Annotate the
dict as dict[str, Any]; the replay path keeps the raw bytes unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: remove 3 tests
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(agent_tools): strip display-format "id:" prefix from model-supplied observation IDs
Observations are presented to agents as [id:xxx], and models sometimes
copy the prefix verbatim despite tool-schema instructions to pass the
bare ID. This silently corrupts source_ids provenance on
create_observations_* (broken links stored in document metadata) and
breaks get_reasoning_chain lookups.
Normalize at both entry points. delete_observations is intentionally
not touched here since #746 already covers it.
Only the "id:" prefix is stripped: document IDs are nanoids whose
alphabet includes "-" and "_", so more aggressive cleanup could mangle
legitimate IDs.
Related to #719.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent_tools): strip whitespace remaining after "id:" prefix removal
Addresses CodeRabbit review: defends against "id: xxx" with a space
after the colon, and matches the docstring, which already promised
surrounding-whitespace stripping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* telemetry: use session and user IDs in langfuse
* test: update old span test
* fix: disable langfuse in unit tests
* fix: add post-loop synthesis span
* refactor: address PR review feedback on langfuse tracing
- Consolidate track_name onto LLMTelemetryContext as the sole home;
remove the honcho_llm_call kwarg and update 4 callers to set it on
telemetry directly. Sentry ai_track now reads telemetry.track_name.
- Decouple escaped-stream self-stamping from run-context exit ordering:
stream_final_response now resets _in_agent_run explicitly around drain.
- Narrow langfuse_agent_step wrap in the tool loop — between-turn
bookkeeping (iteration_callback, choice switch, increment) lifted
outside the span so it scopes only the LLM call + tools.
- Reword test conftest comment to behavior-only language.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: switch langfuse spans to imperative handles
Replaces the context-manager-based langfuse_agent_run/step with imperative
LangfuseAgentRun/Step handles so the run span can outlive the function that
opens it. Streaming responses now own the run handle from construction and
close it after drain, stamping the accumulated streamed text as trace output
(previously blank). Multi-turn generations always stamp provider/model and
step metadata, fixing the regression where only the first turn was annotated.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(llm): record effective prompt-only input on run span
The run-level Langfuse span recorded the raw messages parameter, which is
None for prompt-only calls. Mirror execute_tool_loop's handling and record
the synthesized user message so the trace input isn't blank.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(llm): drop StreamingResponseWithMetadata.__anext__ to prevent span leak
The standalone __anext__ delegated straight to the inner stream, bypassing
the token-folding and Langfuse run-handle close that live only in the
__aiter__ generator. Any caller driving the wrapper via anext() instead of
`async for` would leak the run span and lose final-stream token accounting.
Latent today (all callers use `async for`), removed to close the footgun.
Add tests covering the run-handle drain path: full drain stamps the
accumulated streamed text as the span output and closes once; an abandoned
stream still closes via the finally rather than leaking.
* chore(llm): document intentional empty-body propagate_attributes block
The `with propagate_attributes(...): pass` stamps the active @observe trace
root via the context manager's __enter__ side effect; the empty body reads
as deletable dead code. Add a comment so it isn't removed. Addresses PR review.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(llm): restore api.py types after __anext__ removal
Dropping StreamingResponseWithMetadata.__anext__ made it stop satisfying
the AsyncIterator protocol, breaking the result annotation and the
isinstance narrowing in honcho_llm_call. Widen the tool-less result
annotation to include StreamingResponseWithMetadata and narrow positively
to HonchoLLMCallResponse before reading .content.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: implement read DB and fix queue stale cleanup
* fix: use read_db in internal methods
* fix: mention read db in the CLAUDE.md
* fix: make TRACING checkout hook autocommit-safe; sample cleanup-gate jitter once
The DB.TRACING checkout hook ran `SELECT set_config(...)` at pool checkout,
before the dialect applies the read engine's AUTOCOMMIT isolation level. That
statement autobegins a transaction, and psycopg then refuses to switch the
connection into AUTOCOMMIT ("can't change 'autocommit' now: connection in
transaction status INTRANS"), so every read_only session 500s under TRACING and
the INTRANS connection leaks back to poison later write checkouts. Run the hook
in autocommit and restore the prior mode so it never leaves an open transaction;
set_config(..., is_local=false) is session-scoped and survives the boundary.
Add a regression test (fails without the fix) covering read_only + TRACING.
Also sample the stale-cleanup gate's jittered interval once per attempt instead
of re-rolling it every poll, so the spacing is a fixed deadline per cycle rather
than a random walk (and is testable at non-zero jitter ratios).
* fix: reset request_context in TRACING checkout-hook test
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: add new cloudevents for api routes
* fix: add total input tokens to RepresentationCompletedEvent
* feat(telemetry): inject honcho_version + emitter health metrics
* feat(telemetry): per-LLM-call event with try/finally emission + sampler
Adds LLMCallCompletedEvent (llm.call.completed) — fires once per provider hit
with full cost-attribution context: transport/provider_label, model, token
counts with cache breakdown, finish_reason, outcome (success or error),
is_final_attempt flag, retry/fallback state, duration, tool-call shape,
streaming flag, and agent correlation (run_id + iteration).
- src/telemetry/events/llm.py: new event class + CallPurpose closed enum
(deriver.representation, dialectic.answer, dream.deduction|induction,
summary.short|long). Resource id includes attempt so multi-attempt retries
in one iteration get distinct deterministic ids.
- src/telemetry/events/base.py: BaseEvent._volume_class ClassVar (default
"ground_truth"); the new event opts into "high_volume".
- src/config.py: TelemetrySettings.HIGH_VOLUME_SAMPLE_RATE (default 1.0).
- src/telemetry/emitter.py: deterministic sampler keyed on run_id (so an
entire agent trace is kept or dropped together). Aggregate envelopes
bypass the sampler. Sampled-out events increment the dedicated counter
separate from buffer_full/send_failed drops.
- src/llm/runtime.py: AttemptPlan gains attempt/retry_attempts/is_fallback
so the executor reads retry state without re-deriving it.
- src/llm/types.py: LLMTelemetryContext dataclass carrying workspace,
call_purpose, run_id, iteration, peer fields. Iteration is mutable so
the tool loop can set it per inner call.
- src/llm/executor.py: honcho_llm_call_inner wraps the backend call in
try/finally — emits on success AND on exception, with is_final_attempt
computed from AttemptPlan. Stream path emits a was_stream=True placeholder
(token totals deferred until streaming completion is wired through).
Telemetry failures swallowed.
- src/llm/api.py: threads telemetry kwarg through all 4 signatures into
both honcho_llm_call_inner and execute_tool_loop.
- src/llm/tool_loop.py: _telemetry_for_iteration helper copies the caller
context with iteration set per call — covers both the normal iteration
loop AND the max-iteration synthesis call (iteration N+1).
Tests cover success/error emission, sampler trace-coherence (same run_id →
same decision), volume_class enforcement, unknown call_purpose tolerance,
provider_label inference, and telemetry failure isolation. 378/378 pass.
* feat(telemetry): emit agent.iteration on every LLM response + synthesis
AgentIterationEvent was defined but never emitted on this branch. Phase 2
wires it up in execute_tool_loop so every LLM call inside an agentic loop
produces one event — including the no-tool terminating iteration and the
max-iteration synthesis call — and threads LLMTelemetryContext from dialectic
and dreamer specialists down through honcho_llm_call.
- src/telemetry/events/agent.py: AgentIterationEvent opts into
_volume_class="high_volume" so the Phase 1 sampler throttles it.
- src/llm/tool_loop.py: _emit_agent_iteration() helper fires once per
honcho_llm_call_inner response, BEFORE the no-tool early return so the
terminating iteration is counted. A second emission fires for the
max-iteration synthesis call BEFORE final_response is mutated with
cumulative totals (otherwise the per-iteration counts would double-count).
Emission is defensively skipped when telemetry context lacks run_id /
agent_type / parent_category / workspace_name; emit failures are swallowed.
- src/dreamer/specialists.py: BaseSpecialist.run passes LLMTelemetryContext
with parent_category="dream", agent_type=self.name, observer/observed,
call_purpose=f"dream.{self.name}".
- src/dialectic/core.py: _telemetry_context() builds a shared context for
both answer() and answer_stream(), using self._run_id (always set) +
workspace + observed peer.
Tests cover fresh-copy semantics, per-iteration vs terminating emission,
defensive skip cases, telemetry-failure isolation, and volume_class. 408/408
pass across telemetry + llm + utils + dreamer + dialectic.
* feat(telemetry): agent.tool.call.completed event + ToolResult metadata
Adds the missing generic per-tool-call event so read-only tools (search_*,
get_recent_history, get_observation_context, etc.) and the four existing
state-change tools all produce a telemetry record. Built on a new internal
ToolResult(content, metadata) contract so handlers can surface
search-specific fields (top_k/used_embedding/query_tokens/results_count)
to Phase 3 and create/delete counts to Phase 5's specialist rollups.
- src/telemetry/events/agent.py: AgentToolCallCompletedEvent at v1 with
_volume_class="high_volume". Resource id = {run_id}:{iteration}:{tool_call_seq}
so two calls to the same tool in one iteration don't collide
deterministic ids and get dedup-dropped downstream.
- src/utils/types.py: ToolResult dataclass; two new ContextVars
(_current_tool_call_seq + _last_tool_metadata) so tool_loop and the
execute_tool closure can communicate per-call telemetry without changing
the public Callable[[str, dict], Any] signature.
- src/utils/agent_tools.py: execute_tool times handlers, unwraps ToolResult,
publishes metadata, emits the event. Handlers updated to ToolResult
where useful: create/delete observations, update_peer_card, search_memory,
search_messages. Other handlers continue to return str.
- src/llm/tool_loop.py: set_current_tool_call_seq before each executor call;
read get_last_tool_metadata after and stash on all_tool_calls[i] for
Phase 5 rollups.
Tests cover ToolResult str-likeness, ContextVar round-trip, full-context
emission with search metadata, resource-id disambiguation, defensive skip
cases, telemetry isolation, truncation metadata, volume_class. 420/420 pass.
* feat(telemetry): RepresentationCompletedEvent v2 token breakdown + tool-less truncation
Bulks out the deriver's per-batch telemetry without bumping the event schema
version. New additive fields capture the full token breakdown (queued vs.
extra-context vs. scaffold), the cap configuration (batch_max_tokens,
max_input_tokens, was_flush_enabled), real cap-hit flags, and observer
fanout. `input_tokens` stays unchanged as the queued-message-tokens billing
key Xatu's Stripe meter reads.
The big enabler: src/llm/api.py now actually enforces max_input_tokens on
the tool-less LLM path. Before this, the deriver passed the kwarg but the
path silently dropped it — so the configured cap was advisory and
hit_input_token_cap couldn't be measured. Phase 4 wires truncation through
the same truncate_messages_to_fit helper the tool loop uses and surfaces
input_was_truncated on HonchoLLMCallResponse.
- src/telemetry/events/representation.py: 12 additive fields, schema_version
stays at 2.
- src/llm/types.py: input_was_truncated on HonchoLLMCallResponse.
- src/llm/api.py: tool-less path truncates messages before dispatch, flips
input_was_truncated on the response when clamping occurs. Split into
Literal[True]/Literal[False] branches for typecheck.
- src/deriver/queue_manager.py: QueueBatchResult dataclass replaces the
3-tuple return from get_queue_item_batch; carries hit_batch_token_cap
(computed from cumulative token sum vs cap), was_flush_enabled snapshot,
and batch_max_tokens. Worker loop unpacks + forwards.
- src/deriver/consumer.py: process_representation_batch gains the three
flag kwargs and forwards.
- src/deriver/deriver.py: derives the breakdown fields locally, populates
the new fields on emit, sources hit_input_token_cap from
response.input_was_truncated.
Tests cover schema stability, defaultable fields, input_tokens semantic
preservation, cap-hit flag round-trip, model_dump completeness, and
HonchoLLMCallResponse.input_was_truncated mutability. Existing
test_queue_processing.py tests updated for QueueBatchResult and mock
process_representation_batch signature. 479/479 pass.
* feat(telemetry): DreamRunEvent v2 scheduler reasons + DreamSpecialistEvent v2 rollups
Bumps both dream events to v2 with additive fields. DreamRunEvent gains
scheduler context (threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold / dream_type / enabled_types_count) threaded through the
dream queue payload — the two scheduler gates stay as separate fields rather
than collapsing into one trigger_reason, preserving the WHY-vs-WHEN
semantics. DreamSpecialistEvent gains denormalized rollups
(created_observation_count / deleted_observation_count / peer_card_updated /
search_tool_calls_count) sourced from Phase 3's ToolResult.metadata so the
counts reflect observation truth, not call truth.
- src/telemetry/events/dream.py: schema_version → 2 for both events; new
fields all defaultable so older producers still construct valid events.
- src/utils/queue_payload.py: DreamPayload + create_dream_payload accept
threshold_reason / delay_reason / documents_since_last_dream_at_schedule /
document_threshold.
- src/dreamer/dream_scheduler.py: check_and_schedule_dream computes the two
reasons at decision time and threads them through schedule_dream →
_delayed_dream → execute_dream → enqueue_dream.
- src/deriver/enqueue.py: create_dream_record / enqueue_dream gain the
kwargs and persist on the queue payload.
- src/dreamer/orchestrator.py: process_dream unpacks the payload; run_dream
accepts the kwargs and stamps them on DreamRunEvent.
- src/dreamer/specialists.py: BaseSpecialist.run walks response.tool_calls_made
and sums ToolResult.metadata.created_count / .deleted_count, sets
peer_card_updated, counts search-tool calls by name.
Tests cover schema_version bumps, defaultable Phase 5 fields,
threshold-vs-delay semantics, observation-vs-call-count rollup distinction,
and DreamPayload round-trip. Existing tests updated for the schema bump
and the new enqueue_dream kwargs. 488/488 pass.
* feat(telemetry): AgentToolSummaryCreatedEvent v2 token breakdown
Bumps schema_version to 2 and adds three additive breakdown fields so
analytics can answer "how much of a summary call's cost was the previous-
summary rollup vs. the new messages vs. the scaffold instructions".
- src/telemetry/events/agent.py: previous_summary_tokens, message_tokens,
prompt_scaffold_tokens added with sensible 0 defaults. input_tokens
retains its current semantic (provider-side LLM tokens) — the plan's
proposed `provider_input_tokens` was omitted because input_tokens
already serves that purpose and a duplicate would fork queries.
- src/utils/summarizer.py: emit now populates the three new fields from
values already in scope (messages_tokens, previous_summary_tokens,
prompt_tokens). Hoisted prompt_tokens calculation out of the
is_fallback conditional so both the save-summary path and the emit
share one binding — basedpyright couldn't prove the sibling-scope
binding was safe, and the compute is cheap + idempotent.
Tests cover schema bump, defaultable fields, input_tokens semantic
preservation, first-summary edge case, and breakdown round-trip.
493/493 pass.
* feat(telemetry): embedding.call.completed event + call-purpose ContextVar
Adds the final piece of cost-attribution telemetry: per-embedding-call
events covering every provider hit (single + batch + retry attempts).
Embedding calls are real provider spend that was invisible before this
phase; search-heavy paths (dialectic agentic) can produce more embedding
calls than LLM calls, so the new event participates in the shared
HIGH_VOLUME_SAMPLE_RATE.
- src/telemetry/events/llm.py: EmbeddingCallCompletedEvent at v1 with
_volume_class="high_volume". EmbeddingCallPurpose closed enum
(search_memory / search_messages / create_observations / vector_sync /
summary / message_create). Resource id = run:purpose:provider:model:input_count
so per-iteration calls in one agentic run don't collide.
- src/utils/types.py: _embedding_call_purpose ContextVar plus
@contextmanager wrapper. Nesting-safe via ContextVar.reset(token).
Callers wrap embedding-driving operations in
`with embedding_call_purpose("search_memory"): ...` — no changes to
the embedding client signature.
- src/embedding_client.py: _emit_embedding_call wraps each provider hit
with try/finally so success AND error paths emit. Errors propagate
unchanged. Each retry attempt of _process_batch emits its own event.
Unknown call_purpose slugs drop to None (validation against the enum
happens at emit time, not at context-manager-set time).
- src/utils/agent_tools.py: search_memory / search_messages /
search_messages_temporal / create_observations (batch + fallback) all
tag their embedding calls.
- src/crud/representation.py: save_representation tags with
CREATE_OBSERVATIONS; get_working_representation precompute tags with
SEARCH_MEMORY.
- src/crud/message.py: create_messages batch embed tags with
MESSAGE_CREATE; search_messages/temporal fallback tags with
SEARCH_MESSAGES.
Tests cover event shape, enum closure, ContextVar nesting/exception
cleanup, wrapper success+error emission, unknown-purpose graceful
fallback, telemetry-failure isolation. 550/550 pass across the full
telemetry+llm+utils+dreamer+dialectic+deriver+crud test set.
* chore: fix tests
* fix(telemetry): address review findings on stream events, context propagation, and cap detection
Five findings from a post-Phase-7 review (one resolved by the merge from
main, four addressed here):
- src/llm/executor.py: stream-path LLMCallCompletedEvent now fires AFTER
the stream is set up and drained (or on exception), with real duration
and accurate outcome. Previously the event was emitted before
execute_stream() ran and was always recorded as outcome="success" with
duration_ms=0, which silently masked stream-setup and stream-drain
failures. Wrapping the async generator in try/finally surfaces the real
outcome; token counts stay 0 because we still don't have them at stream
end (aggregate envelopes carry totals).
- src/deriver/deriver.py + src/utils/summarizer.py: deriver and summarizer
LLM calls now thread LLMTelemetryContext into honcho_llm_call. Before
this, the closed CallPurpose enum had DERIVER_REPRESENTATION /
SUMMARY_SHORT / SUMMARY_LONG slugs but those production call sites
didn't actually pass `telemetry=`, so their LLMCallCompletedEvents lost
workspace_name, parent_category, and call_purpose. summarizer threads
workspace_name through _create_and_save_summary → _create_summary →
create_short_summary / create_long_summary.
- src/utils/types.py + src/embedding_client.py: embedding_call_purpose
ctx manager now accepts workspace_name and run_id kwargs, backed by
two new ContextVars. EmbeddingCallCompletedEvent's publisher reads
both via get_embedding_workspace_name / get_embedding_run_id so
embedding events carry workspace and run correlation. All call sites
updated: search_memory / search_messages / search_messages_temporal /
_handle_create_observations_impl pass ctx.workspace_name +
ctx.run_id; create_observations standalone and create_messages pass
workspace_name; RepresentationManager.save_representation and
get_working_representation pass self.workspace_name.
- src/deriver/queue_manager.py: hit_batch_token_cap detection rewritten.
Previously summed kept-rows' token_count and checked against
batch_max_tokens, but the SQL filter `cumulative_token_count <= cap`
guarantees kept rows stay under the cap, so the flag almost never
fired. Now uses two follow-up queries: total token_count across the
included id range + EXISTS check for any session message past the
last-kept id. Both true → cap was actually binding.
(The fifth finding — deriver scaffold-token computation needing
estimate_deriver_prompt_tokens(custom_instructions) — was resolved by
the merge from main; the Phase 4 emit at src/deriver/deriver.py:283
already sources prompt_scaffold_tokens from the wrapped helper.)
567/567 telemetry+llm+utils+dreamer+dialectic+deriver+crud tests pass.
ruff + basedpyright clean.
* chore: ruff linting
* chore: clean AI generated comments references specs
* fix: address coderabbit changes
* fix(telemetry): address remaining PR review findings
Six findings from the PR 637 telemetry review batched into one commit.
- src/llm/executor.py + src/embedding_client.py: asyncio.CancelledError
now surfaces as outcome="cancelled" on both stream and sync paths,
distinct from "error". Client disconnects mid-stream and server
shutdowns are normal control flow and should not feed error-rate
alerting. LLMCallCompletedEvent and EmbeddingCallCompletedEvent
outcome Literal extended; docstrings + tests cover the new state.
- src/utils/types.py + src/llm/tool_loop.py: new iteration_scope()
context manager captures and resets the four per-tool-loop
ContextVars (_current_iteration, _current_tool_call_seq,
_current_provider_tool_call_id, _last_tool_metadata). Applied as a
typed decorator to execute_tool_loop so back-to-back loops in the
same asyncio Task (worker batches, tests) don't observe stale state.
- src/telemetry/events/api.py + src/routers/messages.py:
MessageCreatedEvent schema v1 → v2. Added required last_message_id
(nanoid public_id of the trailing message); get_resource_id now keys
on it instead of message_count, eliminating the collision case where
two same-size batches in the same session+source produced identical
event ids. message_count stays on the body for analytics.
- src/deriver/queue_manager.py: hit_batch_token_cap now computed from
the FINAL post-config-filter batch. Previously the flag used the
pre-filter messages_context[-1].id, which produced false positives
when _resolve_batch_configuration trimmed the trailing queue item —
telemetry reported a cap-hit when the actual returned batch was
short for unrelated reasons. Cap-detection block moved inside the
async with after the filter; no extra DB connection.
- src/config.py + src/telemetry/emitter.py: documented the
HIGH_VOLUME_SAMPLE_RATE orphan trade-off (rate<1.0 keeps aggregates
but drops children, so JOIN ON run_id queries see partial traces).
Behavior unchanged — rate defaults to 1.0.
- src/deriver/deriver.py: WARNING-level invariant logs when
response.input_tokens < messages_tokens (provider tokenization
drift) or prompt_scaffold_tokens <= 0 (estimator silent failure).
Best-effort — telemetry never bleeds into the deriver path
* fix(telemetry): stream retry, embed attempts, truncation, dedup
Address remaining audit findings on the cloudevents PR:
- Stream setup now runs inside the awaited honcho_llm_call_inner so
tenacity's retry wrapper in stream_final_response catches transient
setup failures (rate-limit, auth, network). Previously the returned
generator deferred execute_stream until first iteration — outside
the retry wrapper — crashing the request and bypassing telemetry.
- Embedding _emit_embedding_call gains an is_final_attempt parameter;
_process_batch threads the real retry index so dashboards stop
conflating one-shot, mid-retry, and exhausted-retry calls.
- _truncate_tool_output returns (text, original_chars, was_truncated)
and a new _maybe_truncated_result helper wraps in ToolResult when
truncation happens. Five handlers migrated. AgentToolCallCompletedEvent
fields was_truncated and result_chars_before_truncation are now
populated instead of always None/False.
- execute_tool_loop tracks any_iteration_truncated and stamps
input_was_truncated on the final response (both HonchoLLMCallResponse
and StreamingResponseWithMetadata). Dialectic now reports
hit_input_token_cap correctly.
- GetContextEvent.get_resource_id uses empty-string sentinel instead
of literal "none" so a peer named "none" can't collide with absent.
- generate_event_id folds honcho_version into the deterministic id so
same logical event from different deploys produces distinct ids.
* fix(telemetry): address audit findings across LLM/embed/event paths
Three rounds of telemetry audit findings, grouped by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). Previously the
inner generator deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation populated by _truncate_tool_output via
a new _maybe_truncated_result wrapper; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose context manager accepts parent_category.
- 4 new EmbeddingCallPurpose enum values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Queue batch cap detection (P2.1)
- hit_batch_token_cap keys on the pre-config-filter SQL boundary so the
"kept=900 of 1000 cap, next=300 excluded by cap" case reports True
while still avoiding the config-filter false positive.
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3) — top_k,
used_embedding, embedding_query_count, query_tokens, results_count.
Tests: stream-setup retry, stream-retry attempt sequence, post-stream
output_tokens write-back, is_final_attempt matrix, truncation E2E,
tool-loop hit_input_token_cap propagation, honcho_version in event id,
GetContextEvent disambiguation, queue_items_cleaned round-trip.
* fix(telemetry): address audit findings across LLM/embed/event paths
Four rounds of telemetry audit findings (initial + 3 follow-ups), grouped
by area:
Retry correctness
- Stream LLM setup now runs inside the awaited honcho_llm_call_inner so
tenacity's outer retry catches setup failures (Fix 1). The inner
generator previously deferred execute_stream past the retry wrapper.
- stream_final_response bumps the per-retry attempt index via
dataclasses.replace so emitted events show [1, 2, 3] instead of
[1, 1, 1] (Fix 13).
- Embedding _emit_embedding_call takes is_final_attempt; _process_batch
threads the real retry index (Fix 2).
Token + cost reporting
- HonchoLLMCallResponse.hit_input_token_cap (renamed from
input_was_truncated) uses a token-based rule so single-message
over-cap inputs are correctly flagged — the deriver's prompt-only
path used to silently fly through. Propagated through tool_loop's
per-iteration check (Fix 4) and into RepresentationCompletedEvent.
- DialecticCompletedEvent gains hit_input_token_cap; output_tokens now
folds in the final-stream's cumulative usage via
StreamingResponseWithMetadata.__aiter__ (Fix 7).
Queue batch cap detection
- hit_batch_token_cap previously required total_in_range >= cap, which
produced false negatives whenever the kept range didn't fully exhaust
the budget. Replaced with a pre-config-filter SQL boundary check
(P2.1), then further refined to a queue-item boundary comparison
(Fix 14) so trailing-context trimming doesn't false-negative either.
Event emission completeness
- AgentToolCallCompletedEvent's was_truncated /
result_chars_before_truncation now populated by _truncate_tool_output
via _maybe_truncated_result; 5 handlers migrated (Fix 3).
- DreamSpecialistEvent emits on failure with success=False + new
error_class field, via try/finally (Fix 11). except BaseException
catches cancellations too (Fix 16).
- DeletionCompletedEvent emits on failure paths via try/finally
(Fix 12), and uses ValidationException for unsupported types per
project guideline (Fix 17).
- CleanupStaleItemsCompletedEvent.queue_items_cleaned populated from
deleted_count (Fix 8).
Embedding call attribution (Fix 9)
- embedding_call_purpose accepts parent_category.
- 4 new EmbeddingCallPurpose values: DIALECTIC_PREFETCH,
SESSION_CONTEXT_SEARCH, PREFERENCE_EXTRACTION, GENERIC_DOCUMENT_SEARCH.
- Wrapped previously-unattributed sites: dialectic prefetch, session
context search, preference extraction, conclusions search, vector
sync (×2).
Reconciler no longer holds DB session during embedding (Fix 15)
- _sync_documents and _sync_message_embeddings refactored into
three phases per CLAUDE.md guideline: fetch+detach in a small DB
scope, external embedding call without DB locks, writes in a fresh
short-lived DB scope. New _apply_*_sync helpers; orchestrators
expunge ORM objects before invoking. Vector store upsert + sync_state
updates stay in the apply phase together.
Deterministic event ID + dedup
- generate_event_id folds honcho_version into the hash so cross-deploy
events don't silently collide on ID (Fix 6).
- GetContextEvent resource_id uses empty-string sentinel instead of
"none" so a peer literally named "none" can't collide (Fix 5).
Tool result metadata
- search_messages_temporal returns ToolResult with the same search_meta
shape as search_memory / search_messages (P2.3).
- Dialectic.prefetched_conclusion_count uses Representation.len() so
inductive + contradiction observations count too (Fix 10).
* fix(telemetry): orchestrator emit + review feedback
Three more rounds of audit findings + inline PR review, grouped:
Orchestration / emit reliability
- run_dream wrapped in try/finally so DreamRunEvent always emits, even
on unexpected exceptions including CancelledError (`finally` still
runs while cancellation propagates). Specialist except clauses
broadened from SpecialistExecutionError (never raised in src/) to
Exception so provider/DB/tool failures are recorded with
deduction_success=False / induction_success=False instead of crashing
past the emit.
- BaseSpecialist.run() telemetry state initialization + try/finally
hoisted above the preflight phase (peer lookup, peer-card preload,
create_tool_executor, get_model_config, prompt construction) so
preflight failures emit DreamSpecialistEvent(success=False) instead
of being dropped on the floor.
- Reverted the Round-4 _sync_documents / _sync_message_embeddings
phase split. The split introduced a race: rows were released from
FOR UPDATE SKIP LOCKED before the embed call, allowing two workers
to claim and clobber the same batch. Long-held DB transaction
restored (pre-existing CLAUDE.md violation accepted as a deliberate
trade-off; proper fix requires a claim/in_flight migration tracked
separately).
Schema + naming (PR-internal — none of these have shipped)
- threshold_reason → trigger_reason on DreamRunEvent, DreamPayload, and
every emit/scheduler/router/test call site (~45 src + 21 test lines).
Name now accurately reflects the field's role across "manual",
"surprisal", and "document_threshold" values.
- MessageCreatedEvent reset to schema v1 (was internally bumped to v2
for last_message_id but never shipped at v1 — downstream sees it
for the first time at merge).
- DreamSpecialistEvent gains created_counts_by_level /
deleted_counts_by_level: dict[str, int] keyed on the closed
level taxonomy. Per-tool-call events use list[str] (≤10 items),
but specialist runs aggregate 20+ — dict keeps emissions compact.
- QueueBatchResult marked frozen=True.
Per-call embedding attribution
- Agent tool embedding_call_purpose wraps for search_memory,
search_messages, search_messages_temporal, create_observations now
driven embedding cost rolls up under the right workflow.
- create_observations() signature gains parent_category kwarg
(mirrors existing run_id pattern).
Manual dream scheduling
- Manual /schedule_dream route now passes trigger_reason="manual" and
delay_reason="immediate". Previously both arrived as null in
DreamRunEvent, breaking analytics joins.
Queue-batch SQL perf
- next_exists_check folded into the main CTE query via
bool_or(cumulative_token_count > batch_max_tokens) OVER () in a
nested subquery. Cap detection is now one roundtrip per batch
instead of two.
Code/doc cleanup
- representation.py docstring uses generic "downstream metering key"
language (was "Xatu's Stripe meter"). bench runner --base-url help
uses a generic example host (was "groudon.fly.dev"). Public-facing
code/docs shouldn't reference internal service names.
Tests added for: orchestrator failure-path DreamRunEvent emission,
specialists preflight try/finally coverage, manual-dream
trigger_reason/delay_reason round-trip, dict-rollup accumulation across
multiple tool calls in a specialist run, CTE-fold one-roundtrip
behavior. Full Python suite passes (1236).
* fix(telemetry): correctness + attribution + emitter robustness
- Dreamer iteration count: read response.iterations directly so
one-shot runs no longer report iterations=0 and tool-using runs
include the terminal/synthesis LLM call.
- RepresentationCompletedEvent.observer_count counts successful
saves, not attempts.
- search_memory empty-memory fallback reports the snippet count when
message context is returned (was always 0).
- Wire parent_category through every embedding emit path: message
create (api), save_representation (representation), per-observation
fallback (caller-supplied), and the peer/session context routes
(api). get_working_representation accepts parent_category and
embedding_purpose so the internal fallback embed lands in the same
analytics bucket as the route-level precompute even when the
precompute is suppressed.
- BatchItem carries token_count so _process_batch reuses chunk-prep
counts instead of re-encoding every chunk for the telemetry proxy.
- Drop vestigial EmbeddingCallCompletedEvent.batch_size (always ==
input_count).
- Emitter: release the lock during HTTP send so a failing endpoint's
retry+backoff (~36s worst case) doesn't block other flushers;
edge-trigger the 80%-capacity warning so sustained backpressure
doesn't flood logs; defer event_id generation past the high-volume
sampler for events with run_id so sampled-out children don't pay
the sha256; harden emit() against sync callers with no running
loop; track threshold-flush tasks so shutdown() drains in-flight
sends before closing the HTTP client.
* fix(telemetry): tool cancellation emit, nanoid run_ids, version unification
- execute_tool: wrap post-work in finally so AgentToolCallCompletedEvent
fires on CancelledError; explicit handler sets is_error/result_str
before re-raising.
- run_id: replace str(uuid.uuid4())[:8] with generate_nanoid() across
dialectic/dreamer/specialists; matches project-wide nanoid convention.
- Bump _schema_version on events touched by run_id widening:
DialecticCompletedEvent v1→v2 (also covers hit_input_token_cap field),
AgentIterationEvent v1→v2, AgentToolConclusionsCreatedEvent v1→v2,
AgentToolConclusionsDeletedEvent v2→v3, AgentToolPeerCardUpdatedEvent
v1→v2.
- Unify honcho_version: single HONCHO_VERSION constant in src/_version.py
read from pyproject.toml (importlib.metadata fallback). Drop
TELEMETRY.HONCHO_VERSION setting. Use the constant for the FastAPI app
version (no more hardcoded "3.0.6") and for emitter body injection.
- Delete 17 tautological per-event test_schema_version methods; the
parametrized contract test still enforces version >= 1 across all events.
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix(deriver): ignore blank observations before embedding
* Address PR review on observation normalization
* Harden mock await arg access in tests
* Unify blank observation filtering across tool paths
* Move soft-delete query test back to fixture class
* fix: Add JSON repair for truncated LLM responses across all providers and Gemini thinking budget support
LengthFinishReasonError from OpenAI-compatible providers (custom, openai, groq) was crashing the deriver
with 14k+ occurrences in production. The vLLM path already had repair logic but it was gated on
provider=="vllm", unreachable when routing through litellm as a custom provider.
- Extract shared _repair_response_model_json() helper for all providers
- Catch LengthFinishReasonError in OpenAI/custom parse() path and repair truncated JSON
- Add repair fallback to Anthropic and Gemini response_model paths
- Add repair fallback to Groq response_model path
- Pass thinking_budget_tokens to Gemini 2.5 models via thinking_config
- Add 14 tests covering repair paths for all providers and Gemini thinking budget
Fixes HONCHO-YC
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: live llm integration tests
* feat: Consistent Model Config Protocol
* fix: migrate the remaining app callers off the legacy llm_settings path
* fix: Docs and regression tests
* fix: refactor llm runtime path to model-config-only API
* fix: refactor config to nested model-config source of truth
* fix: refactor llm streaming and tool dispatch through backends
* fix: cut over llm config to nested model_config only
* fix: collapse vllm and custom into openai_compatible transport
* feat: refactor llm config to explicit transports and bare model ids
* feat: (embed) Add configurability for embedding model
* fix: tests for embedding provider
* fix: Address Review Comments
* fix: (llm) remove Groq backend and per-vendor base URLs
* chore: move llm tests
* fix: (llm) address review findings — config regressions, backend bugs, dead code
* fix: address backend end silly errors
* chore: (docs) update configuration and self-hosting guides
* chore: fix tests
* fix: address code rabbit comments
* fix: add validation to the dream settings
* fix: further address code rabbit comments
* fix: Address Code Rabbit Comments
* fix: Another round of code rabbit
* fix: Address Code Rabbit Nits
* fix: tests
* refactor: rename thinking validator to reflect transport scope
_validate_anthropic_thinking_minimum only enforces the >=1024 rule for
Anthropic and no-ops for other transports, so the name was misleading
now that it's shared across ConfiguredModelSettings, FallbackModelSettings,
and ModelConfig. Renamed to _validate_thinking_constraints with a docstring
clarifying per-transport behavior. No logic change.
* fix(config): drop transport-specific thinking params when env override changes transport
_fill_defaults_for_nested_field previously preserved the default MODEL_CONFIG's
thinking_budget_tokens/thinking_effort across a transport override. This leaked
Gemini-family defaults (e.g. thinking_budget_tokens=1024) into OpenAI-transport
overrides, and the OpenAI backend then correctly rejected the unsupported param
at call time (OpenAI uses reasoning.effort, not a token budget).
The helper now strips thinking_budget_tokens and thinking_effort from the
default dict when the env override supplies a transport different from the
default's. Explicit thinking params in the override are preserved.
* fix(config): apply thinking-param strip to dialectic level merge too
DialecticSettings._merge_level_defaults does its own inline MODEL_CONFIG
merge (parallel to _fill_defaults_for_nested_field), so the previous fix
missed dialectic-level overrides. E.g. flipping
DIALECTIC_LEVELS__minimal__MODEL_CONFIG__TRANSPORT from gemini (default)
to openai still leaked the default thinking_budget_tokens=0 into the
openai config, which the OpenAI backend then rejected at call time.
The level-merge path now applies the same 'strip transport-specific
thinking params when transport changes' rule as the generic helper.
Added a regression test exercising the merge validator directly.
* refactor(llm): wire ModelConfig knobs through, prune clients.py migration leftovers
Three connected fixes to finish carving the LLM stack out of src/utils/clients.py
and into src/llm/:
1. Propagate ModelConfig tuning knobs into backend calls.
honcho_llm_call_inner built extra_params from only {json_mode, verbosity},
silently dropping top_p, top_k, frequency_penalty, presence_penalty, seed,
and operator-supplied provider_params from any ModelConfig. Thread the
selected config through ProviderSelection and merge
build_config_extra_params(selected_config) into extra_params; per-call
kwargs still win over provider_params defaults. Makes
_build_config_extra_params public as build_config_extra_params so
clients.py and request_builder.py share one translation. Adds
TestModelConfigExtraParamsPropagation covering OpenAI/Anthropic knob
propagation, provider_params passthrough, and per-call override
precedence.
2. Drop dead extract_openai_* duplicates in clients.py.
extract_openai_reasoning_content, extract_openai_reasoning_details, and
extract_openai_cache_tokens had no callers outside their own definitions
— the live implementations live in src/llm/backends/openai.py. -103
lines from clients.py.
3. Unify on ModelTransport, delete SupportedProviders.
The "google" vs "gemini" split forced a _provider_for_model_config
translation shim in two places. Replace all SupportedProviders usages
with ModelTransport, rename CLIENTS["google"] → CLIENTS["gemini"],
update provider branches + LLMError labels + reasoning-trace entries
accordingly. Trace JSONL now writes "provider": "gemini" instead of
"google" — consistent with the broader env-var rename cutover.
Also tidies up pre-existing basedpyright findings in tests/llm/test_model_config.py
(pydantic before-validator dict inputs + descriptor-proxy call).
ruff: clean. basedpyright: 0 errors, 0 warnings. Tests: 153/153 pass across
tests/utils/test_clients.py, tests/utils/test_length_finish_reason.py,
tests/llm/, tests/dialectic/, tests/deriver/.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): finish the src/utils/clients.py → src/llm/ migration
honcho_llm_call_inner now delegates to request_builder.execute_completion
and execute_stream instead of re-implementing backend call scaffolding
inline. The new _effective_config_for_call helper carries per-call kwargs
(temperature, stop_seqs, thinking_budget_tokens, reasoning_effort) onto
the selected ModelConfig — or synthesizes a minimal config for the
test-only callers that pass provider+model directly. max_output_tokens
is zeroed on the effective config to preserve the current
"per-call max_tokens wins" semantic; honoring ModelConfig.max_output_tokens
is a separable correctness concern.
Side effect of routing through the new path: ConfiguredModelSettings'
thinking_budget_tokens validator now fires on synthesized configs.
test_anthropic_thinking_budget was asserting that a sub-1024 budget
propagated to Anthropic — bumped to 1024 to match what Anthropic actually
accepts.
Unified client construction. Promoted the cached client factories in
src/llm/__init__.py (get_anthropic_client, get_openai_client,
get_gemini_client, get_{anthropic,openai,gemini}_override_client) to
public API and added them to __all__. Promoted
credentials._default_transport_api_key → default_transport_api_key.
Deleted the duplicate _build_client and _default_credentials_for_provider
from clients.py; _client_for_model_config now falls through to the
public factories. CLIENTS dict and _get_backend_for_provider stay as the
mockable seam for the ~50 patch.dict(CLIENTS, {...}) test call sites.
Wired operator-configurable Gemini cached-content reuse end-to-end.
PromptCachePolicy moved from src/llm/caching.py into src/config.py so
ModelConfig can reference it as a field without a circular import;
caching.py re-exports the name for existing imports. Added
cache_policy: PromptCachePolicy | None on ConfiguredModelSettings,
FallbackModelSettings, ResolvedFallbackConfig, and ModelConfig.
resolve_model_config, _resolve_fallback_config, and
_select_model_config_for_attempt copy the field through.
honcho_llm_call_inner passes effective_config.cache_policy into
execute_completion / execute_stream, so operators opt in via
e.g. DERIVER_MODEL_CONFIG__CACHE_POLICY__MODE=gemini_cached_content
and the selection actually fires instead of sitting on a dead path.
New regression test test_cache_policy_reaches_gemini_backend asserts the
PromptCachePolicy object reaches the Gemini backend's extra_params.
ruff + basedpyright: clean. Tests: 154/154 pass across
tests/utils/test_clients.py, tests/utils/test_length_finish_reason.py,
tests/llm/, tests/dialectic/, tests/deriver/.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): move all LLM orchestration into src/llm/ and delete clients.py
The 1624-line src/utils/clients.py has been carved up into focused modules
under src/llm/ and deleted. There is now one golden path for LLM
orchestration and no dual entrypoint.
New module layout:
src/llm/
__init__.py thin stable re-export surface
api.py public honcho_llm_call with retry + fallback + tool
loop delegation
executor.py honcho_llm_call_inner (single-call executor); bridges
to request_builder.execute_completion / execute_stream
tool_loop.py execute_tool_loop + stream_final_response, plus
assistant-tool-message and tool-result formatting
runtime.py AttemptPlan dataclass (replaces the loose
ProviderSelection NamedTuple), effective_config_for_call,
plan_attempt, per-retry temperature bump, attempt
ContextVar
registry.py single owner of CLIENTS dict + cached default and
override SDK-client factories + backend/history-adapter
selection + high-level get_backend(config)
conversation.py count_message_tokens, tool-aware message grouping,
truncate_messages_to_fit
types.py HonchoLLMCallResponse, HonchoLLMCallStreamChunk,
StreamingResponseWithMetadata, IterationData,
IterationCallback, ReasoningEffortType, VerbosityType,
ProviderClient
request_builder.py low-level request assembly (ModelConfig → backend
complete/stream); no longer owns credential resolution
credentials.py default_transport_api_key, resolve_credentials
caching.py gemini_cache_store; re-exports PromptCachePolicy
from src.config
backend.py Protocol + normalized result types
history_adapters.py provider-specific assistant/tool message shapes
structured_output.py
backends/ AnthropicBackend, OpenAIBackend, GeminiBackend
handle_streaming_response had no production callers; it is deleted. The
three tests that used it now drive honcho_llm_call_inner(stream=True,
client_override=...) directly, which exercises the same code path the
public API uses.
Dead credential passthrough removed. The ProviderBackend Protocol and
all three concrete backends no longer accept api_key / api_base — those
are baked into the underlying SDK client at registry construction time
and were being del'd everywhere they appeared. request_builder also
stops resolving and forwarding them.
Client construction is unified. The cached default-client factories
(get_anthropic_client, get_openai_client, get_gemini_client) and override
factories (get_*_override_client) are promoted to public API; the
module-level CLIENTS dict populates from them and remains the
patch.dict(CLIENTS, {...}) mocking seam tests rely on. Old duplicate
helpers (_build_client, _default_credentials_for_provider) are gone.
default_transport_api_key is promoted to public.
Application imports now come from src.llm (dreamer, dialectic, deriver,
summarizer, telemetry-adjacent tests). No code imports from
src.utils.clients anywhere in the repo.
ruff: clean. basedpyright: 0 errors, 0 warnings. Tests: 1013/1013 pass
across the entire non-infra test suite (excluding tests/unified,
tests/bench, tests/live_llm, tests/alembic).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(llm): sanitize tool schemas for Gemini's function_declarations validator
Gemini's native-transport function-declarations validator accepts a narrow
subset of JSON-Schema / OpenAPI: type, format, description, nullable, enum,
properties, required, items, minItems, maxItems, minimum, maximum, title.
Anything else — additionalProperties, allOf, if/then/else, $ref, anyOf,
oneOf, $defs, patternProperties — triggers an INVALID_ARGUMENT 400 at call
time.
Our agent tool schemas in src/utils/agent_tools.py use several of those
(additionalProperties: false, allOf + if/then conditionals) because they
were authored for OpenAI strict-mode + Anthropic, which need the richer
vocabulary. GeminiBackend._convert_tools was passing them straight through.
Add _sanitize_schema(): walks the parameters tree and drops unsupported
keywords while preserving semantics for the keywords that hold user data
(properties maps field-name → sub-schema; required / enum are lists of
literals; items is a single sub-schema). Other backends are untouched and
continue to receive the full strict schemas.
Regression tests:
- test_gemini_sanitize_schema_strips_unsupported_keywords: confirms
additionalProperties, allOf + if/then, and $defs are stripped at nested
levels while legitimate fields survive.
- test_gemini_convert_tools_sanitizes_parameters_schema: end-to-end
_convert_tools output has no forbidden keys.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: fix tool calling syntax for gemini
* refactor(llm): normalize defaults, widen OpenAI reasoning-model routing
* chore: fix test
* fix(llm): address post-migration review feedback
* fix(llm): gemini robustness + dreamer specialist ergonomics
* chore: addres review comments
* chore: (docs) unrelease changelog addition
* chore: (docs) merge commit changes
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Erosika <eri@plasticlabs.ai>
* fix: further remove extraneous transactions
* fix: (search) use 2 phase function to reduce un-needed transaction
* fix: refactor agent search to perform external operations before making a transaction
* fix: reduce scope of queue manager transaction
* fix: (bench) add concurrency to test bench
* fix: address review findings for search dedup, webhook idempotency, and bench throttling
* Fix Leakage in non-session-scoped chat call (#526)
* fix: (search) reduce scope for peer based searches
* fix: tests
* fix: (test) address coderabbit comment
* fix: drop db param from deliver_webhook
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: dialectic held connection
* fix: (agent) pre-compute embeddings for agent tools
* fix: (tests) refactor tests to use smaller test db connections
* fix: Embedding client to branch depending on vector store
* fix: reflect dedup-skipped observations in created counts and isolate DB sessions in extract_preferences
* fix: (tests) update tests to match changes
* fix: expunge docs + don't pass in db to query_documents
---------
Co-authored-by: Rajat Ahuja <rahuja445@gmail.com>
* fix: Add bounds to gemini client
* fix: Prevent empty summaries from being saved to DB (HONCHO-M7)
Raise LLMError on blocked Gemini responses (SAFETY, RECITATION, etc.)
so retry/backup-provider logic triggers. Treat empty LLM responses in
the summarizer as fallback instead of persisting empty strings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: Summary Eval via Locomo
* fix: Code Rabbit Comments
* fix: Code Rabbit Comments
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: Use savepoints to prevent race condition in get_or_create chain
Replace commit()/rollback() with begin_nested() savepoints in
get_or_create_workspace, get_or_create_peers, and get_or_create_session
so that an IntegrityError rollback in a nested call doesn't undo flushed
work from the caller. Moves transaction commit responsibility to the
outermost caller and defers cache operations to post-commit callbacks
on GetOrCreateResult.
Closes DEV-1321
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: Add missing commit and post_commit in chat endpoint
The chat() endpoint in routers/peers.py called get_or_create_peers()
but discarded the result without committing or invoking post_commit().
This meant new peers created lazily via SDK chat calls were never
persisted, and cache invalidation was skipped.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat: implement async workspace deletion with active session checks
- Updated the DELETE /workspaces/:id endpoint to return 202 Accepted, indicating that the deletion request is processed in the background.
- Added a check for active sessions before allowing workspace deletion, raising a ConflictException if any exist.
- Updated related tests to ensure proper handling of active sessions during workspace deletion.
* fix: Address review issues
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* refactor: enhance specialist exploration and observation processes
- Updated the orchestration logic to allow specialists to explore freely with optional hints from high-surprisal observations.
- Removed predefined probing questions, enabling a more dynamic approach to observation gathering.
- Adjusted the deduction and induction specialists to utilize the current peer card context and exploration hints in their prompts.
- Improved documentation within the code to clarify the roles and responsibilities of specialists in the observation process.
* feat: add thorough dream test
* refactor: standardize hint terminology and dedup peer card context
* feat: make session_name nullable for documents and update related SDKs
- Introduced a migration to make the `session_name` column in the documents table nullable, allowing for sessionless dreams.
- Updated Python and TypeScript SDKs to reflect the optional nature of `session_id` in conclusion creation and related methods.
- Enhanced tests to cover scenarios for creating conclusions without a session ID, ensuring proper handling of sessionless conclusions.
- Adjusted documentation and type definitions to clarify the optional session context in various components.
* chore: add migration test
* fix: ensure orphaned sessions exist during downgrade for nullable session_name migration
* chore: 3.0 honcho and 2.0 sdks changelog
fix: use PeerContextResponse in peer.ts
* chore: move docs to /v3/, build SDKs
* chore: code review
* feat: [WIP] migrate away from stainless in typescript sdk
* chore: move api from /v2/ to /v3/
* feat: no-stainless typescript with real tests
* feat: migrate python sdk off of stainless
* feat: clean typescript sdk
* chore: add tests for ts http client
* fix: rewrite entire python sdk in new format, update typescript sdk to use `configuration` not `config` for consistency with API
* fix: clean up SDKs, synchronize
* chore: update sdk examples
* chore: update OpenAPI documentation and SDK examples to reflect changes
* fix: better test
* fix: install deps in test runner, improve robustness of streaming in sdk, coderabbit nits
* fix: standardize around camelCase in TS SDK
* refactor: update configuration handling in SDKs to use typed models for workspace, session, and peer configurations
* docs: clarify queue status usage and remove polling methods from SDKs
add claude skills for migrations
* chore: fix links in docs
* feat: add deriver flush mode to bypass batch token threshold
- Introduced `is_deriver_flush_enabled` function to check if flush mode is active.
- Updated `QueueManager` to conditionally apply batch token thresholds based on flush mode.
- Enhanced `UnifiedTestExecutor` to enable flush mode via Redis.
- Added `flush` parameter to test cases to facilitate testing of flush mode behavior.
- Updated various test cases to utilize the new flush functionality.
* feat: implement schedule_dream functionality in SDKs, use in unified test runner
- Added `schedule_dream` method to both Python and TypeScript SDKs for scheduling dream tasks.
- Updated HTTP routes to include endpoint for scheduling dreams.
- Enhanced test runner to utilize the new `schedule_dream` method for scheduling actions.
- Updated TypeScript client to support the new scheduling functionality with appropriate parameters.
* feat: update single deriver task to support multiple observers
- Changed the `observer` parameter to `observers` as a list in multiple functions across the deriver module.
- Updated the processing logic to handle multiple observers for representation tasks.
- Adjusted related payload and queue management functions to accommodate the new observers structure.
- Modified tests to reflect changes in the representation task handling and ensure proper functionality.
* refactor: update enqueue tests to support deduplication of queue items with multiple observers
- Modified tests in `test_enqueue.py` to reflect changes in the queue item structure, where each message now results in a single queue item containing a list of observers.
- Updated assertions to validate that the `observers` field correctly includes all relevant peers, ensuring proper functionality of the deduplication logic.
- Removed redundant payload matching logic to streamline test cases and improve clarity.
* fix: add backwards compatibility for representation work unit keys and payload observers
* feat: update dialectic configuration and introduce cost calculator
- Adjusted LLM and dialectic settings in `.env.template`, `config.toml.example`, and `src/config.py` to reduce maximum tool output characters and session history tokens for cost efficiency.
- Implemented a new `dialectic_cost_calculator.py` script to estimate costs based on reasoning levels and model pricing.
- Enhanced `DialecticAgent` to utilize minimal tools and adjusted output token settings based on reasoning level to optimize performance and reduce costs.
* feat: add reasoning level to chat input in unified test runner
- Enhanced the `UnifiedTestExecutor` to include a `reasoning_level` parameter in the chat method call.
- Updated the `QueryAction` model to support the new `reasoning_level` attribute, allowing for more nuanced chat interactions.
* feat: run deriver once for multiple observers (#335)
* feat: update single deriver task to support multiple observers
- Changed the `observer` parameter to `observers` as a list in multiple functions across the deriver module.
- Updated the processing logic to handle multiple observers for representation tasks.
- Adjusted related payload and queue management functions to accommodate the new observers structure.
- Modified tests to reflect changes in the representation task handling and ensure proper functionality.
* refactor: update enqueue tests to support deduplication of queue items with multiple observers
- Modified tests in `test_enqueue.py` to reflect changes in the queue item structure, where each message now results in a single queue item containing a list of observers.
- Updated assertions to validate that the `observers` field correctly includes all relevant peers, ensuring proper functionality of the deduplication logic.
- Removed redundant payload matching logic to streamline test cases and improve clarity.
* fix: add backwards compatibility for representation work unit keys and payload observers
* feat: refactor benchmark runners to share common functionality
- Introduced a new `runner_common.py` module containing shared utilities for benchmark test runners, including common argument parsing, client creation, and queue management.
- Updated `BEAMRunner`, `LoCoMoRunner`, and `LongMemEvalRunner` to inherit from `RunnerMixin`, leveraging shared functionality for metrics collection and logging.
- Added `reasoning_level` and `redis_url` parameters to runner constructors for enhanced configuration.
- Streamlined argument parsing by utilizing `add_common_arguments` for shared command-line options across all runners.
* fix: update last_user_message handling to use message content instead of ID
* fix: standardize config vs configuration
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>