Commit Graph

477 Commits

Author SHA1 Message Date
adavyas bd801f4964 Fix basedpyright regressions for custom instructions 2026-04-07 12:33:34 -04:00
adavyas c6ca57b9d8 Fix schema validation typing test coverage 2026-04-07 12:25:20 -04:00
adavyas de82b80227 Require explicit deriver custom instruction token cap 2026-04-07 12:11:26 -04:00
adavyas cd1c35c720 feat: cap deriver custom instruction tokens 2026-04-07 10:21:09 -04:00
adavyas 1b27bff625 Align custom instruction tests with split deriver prompts 2026-04-06 15:55:08 -04:00
adavyas f763f32b04 Fix custom instruction test typing 2026-04-06 15:55:08 -04:00
adavyas 6a746f1de4 Implement deriver custom instructions 2026-04-06 15:55:08 -04:00
adavyas 8a42a9b5b9 Stabilize prompt cache layout tests 2026-04-06 15:39:39 -04:00
adavyas d263e49221 Refine deriver prompt wording 2026-04-06 15:32:57 -04:00
adavyas db94ac050f Fix empty search query validation 2026-03-28 20:10:47 -07:00
adavyas 131905d14a Merge branch 'main' into codex/fix/prefix-based-cache-optim
# Conflicts:
#	src/schemas/api.py
2026-03-28 19:59:43 -07:00
adavyas 15d0938a9b Refine deriver prompt examples 2026-03-28 19:41:05 -07:00
adavyas aa326c8870 Harden schema sanitization and prompt cache tests 2026-03-27 15:08:59 -07:00
adavyas 2c4d88c316 Tune dialectic cache layout for Gemini 2026-03-27 07:45:56 -07:00
adavyas 3ea34064e1 Fix Gemini streaming and deriver cacheability 2026-03-25 00:00:34 -07:00
Viktor Szépe b3c1917a38
Fix typos (#440) 2026-03-22 15:52:55 -04:00
ajspig 92914abdeb
Abigail/dev 1436 (#438)
* docs: comprehensive overview guide with all recent integrations

- Add all recent integrations: Gmail, Granola, OpenClaw, Agent Zero, Hermes
- Organize by use case: Quick Start, Agent Frameworks, Platform Integrations, Community Agents, Communication Platforms, Data Import
- Include feature comparison matrix with support levels
- Add clear categorization by setup time and complexity
- Structure follows OpenClaw documentation format with clear sections

* docs: update overview guide to match dev-1372-new format

- Use simpler, cleaner organization with focused sections
- Group by: AI Assistants, Platform Connectors, Agent Frameworks, Migrations
- Match icon and formatting style from dev-1372-new branch
- Include all recent integrations: Gmail, Granola, Hermes, Agent Zero, OpenClaw

* docs: clean up Platform Connectors section description

Remove redundant description text for cleaner presentation
2026-03-22 15:50:01 -04:00
Eri Barrett 254510db63
Merge pull request #427 from plastic-labs/docs/hermes-honcho-guide
docs: add Hermes Agent Honcho integration guide
2026-03-20 17:19:06 -04:00
Erosika d8a01230ae docs: delete old community hermes stub 2026-03-19 18:06:06 -04:00
Erosika e2e309148b docs: move Hermes from community to integrations in sidebar nav 2026-03-19 17:43:48 -04:00
Erosika c9de186b80 docs: update Hermes guide description 2026-03-19 17:42:49 -04:00
Erosika 65e68fae55 docs: trim Hermes guide, remove CLI/gateway mode section 2026-03-19 17:42:49 -04:00
Rajat Ahuja 1cbcbc0263
fix: populate test harness DB config from docker compose (#435) 2026-03-19 17:39:57 -04:00
LRRuan 2097c2cbcf
fix(files): handle empty json uploads safely (#434)
* fix(files): handle empty json uploads safely

* fix(files): normalize invalid json upload errors

* fix(files): restore file processing error import

---------

Co-authored-by: LRRuan <lrruan@users.noreply.github.com>
2026-03-18 18:36:34 -04:00
Vineeth Voruganti 09a980c2fb
Sanitization and Memory Bug Fixes (#419)
* fix: use WeakValueDictionary for _observation_locks to prevent memory leak

* fix: harden input sanitization across API surface (DEV-1400)

- Parameterize SQL in set_config calls to prevent injection via request context
- Strip NUL bytes from string inputs (message content, queries, peer cards)
- Add JSONB metadata validation (100 key limit, 5 depth limit)
- Add filter recursion depth limit (max 5) to prevent stack overflow
- Update changelogs with unreleased entries

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: Refactor Schemas into separate files

* fix: Code Rabbit Comments

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-18 15:01:52 -04:00
Rajat Ahuja 122ebfe998
fix: add observability to docker compose + get docker compose into a usable state (#429)
* fix: add observability to docker compose + get docker compose into a usable state

* fix: init.sql location

* fix(docker): use built venv at runtime

---------

Co-authored-by: adavyas <adavyasharma@gmail.com>
2026-03-18 13:16:22 -04:00
ajspig db6543acf8
Adding granola integration (#417)
* feat: adding python granola example

* docs: adding granola documentation and cleaning up the script

* chore: addressing pr comments

* chore: nav update

* chore: cr updates

* fix: sig restructure and simplification

* docs: updating for clarity

* chore: cr

* fix: cleaning up script

* feat: adding gmail guide

* refactor: replace globals with param passing, add Granola rate limit retry, surface unparseable transcripts

* chore: cr

* fix: adding example script and added more of a natural flow/progression

* fix: final fix
2026-03-18 13:08:20 -04:00
adavyas 6dc080aeaf Use Gemini function names for tool responses 2026-03-14 17:01:20 -07:00
adavyas d2ce88737d Preserve Gemini tool context during provider fallback 2026-03-14 16:49:54 -07:00
adavyas 56924eb956 Fix remaining prefix cache review issues 2026-03-14 16:29:56 -07:00
adavyas 4ceec9f778 Address cache probe and Gemini review feedback 2026-03-14 14:44:08 -07:00
adavyas 1b32a8a75c Harden cache probe namespaces and docstrings 2026-03-14 14:22:05 -07:00
adavyas eebcbab170 Fix basedpyright warnings in cacheable system blocks 2026-03-14 13:24:48 -07:00
adavyas 482bb6cb16 Fix summarizer cache test typing 2026-03-14 13:12:52 -07:00
adavyas e5841882ad Track Gemini cache token metrics 2026-03-14 12:59:55 -07:00
adavyas e86cbac75b Improve prefix cache eval harness 2026-03-13 22:33:23 -07:00
adavyas 1b99469335 Optimize prompt prefix caching 2026-03-13 17:36:06 -07:00
Vineeth Voruganti 24f94f3ff8
chore: (docs) Add information on the peer card (#418) 2026-03-09 21:52:37 -04:00
ajspig f61d58f56e
Abigail/dev 1406 (#416)
* docs: replace Chatbots section with Tutorials, move Reachy Mini into it

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: restructure nav — move modeling-data to Core Concepts, file-uploads to Advanced, merge Migrations into Tutorials

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: final moving

* docs: minor changes

* docs: adding community integrations

* chore: cr updates

* fix: Rename patterns page and nitpicks on organization

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
2026-03-06 23:21:28 -05:00
Vineeth Voruganti 2f895efb7f
chore: (docs) update claude code plugin copy (#414)
* chore: (docs) update claude code plugin copy

* docs: adding reachy-mini

* docs: adding openclaw multi-agent setup

* docs: updating

* chore: cr

---------

Co-authored-by: ajspig <dragon@monstercode.com>
2026-03-05 10:30:01 -05:00
ajspig ee1ffada21
docs: updating to match recent updates (#413) 2026-02-27 12:22:46 -05:00
Vineeth Voruganti f686205167
fix: Update Changelogs and OpenAPI Docs (#412) 2026-02-25 22:16:44 -05:00
Vineeth Voruganti 10ef7b96a8
Add Stricter limits to Summary & Peer Card (#400)
* fix: Add bounds to gemini client

* fix: Prevent empty summaries from being saved to DB (HONCHO-M7)

Raise LLMError on blocked Gemini responses (SAFETY, RECITATION, etc.)
so retry/backup-provider logic triggers. Treat empty LLM responses in
the summarizer as fallback instead of persisting empty strings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: Summary Eval via Locomo

* fix: Code Rabbit Comments

* fix: Code Rabbit Comments

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 15:08:09 -05:00
Vineeth Voruganti 5a42f9b3cd
Use savepoints to prevent race condition in get_or_create chain (#399)
* fix: Use savepoints to prevent race condition in get_or_create chain

Replace commit()/rollback() with begin_nested() savepoints in
get_or_create_workspace, get_or_create_peers, and get_or_create_session
so that an IntegrityError rollback in a nested call doesn't undo flushed
work from the caller. Moves transaction commit responsibility to the
outermost caller and defers cache operations to post-commit callbacks
on GetOrCreateResult.

Closes DEV-1321

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: Add missing commit and post_commit in chat endpoint

The chat() endpoint in routers/peers.py called get_or_create_peers()
but discarded the result without committing or invoking post_commit().
This meant new peers created lazily via SDK chat calls were never
persisted, and cache invalidation was skipped.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 23:28:04 -05:00
Vineeth Voruganti 3b1e964372
Modify Queue Status to only relevant tasks (#398)
* feat: Update Honcho system benchmarks

* fix: Align with runner common functions and fixed basedpyright issues

* fix: Address coderabbit issues

* chore: Address Ruff Errors

* fix: Queue Status to remove unhelpful info

* chore: (docs) Add docs for dreaming and buffering

* chore: Remove claude workflow

---------

Co-authored-by: 3un01a <3un01a@plasticlabs.ai>
2026-02-23 22:54:21 -05:00
3un01a 2631209256
feat: Update Honcho system benchmarks (#393)
* feat: Update Honcho system benchmarks

* fix: Align with runner common functions and fixed basedpyright issues

* fix: Address coderabbit issues

* chore: Address Ruff Errors

---------

Co-authored-by: 3un01a <3un01a@plasticlabs.ai>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
2026-02-23 21:30:30 -05:00
3un01a eba9279af2
Oolong Benchmark (#323)
* (feat) Add Oolong Benchmarks

* (fix) Address issues to fix basedpyright and coderabbit comments

* (fix) Address basedpyrwright additional warnings

* (fix) Address additional coderabbit issues

* (fix) Replace huggingface data loading to local filesystem-based

* (fix) Address coderabbit issues regarding data paths

* fix: Align with test harness conventions

* fix: Code Review Comments

* fix: stream data rather than load all at once

---------

Co-authored-by: 3un01a <3un01a@plasticlabs.ai>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
2026-02-23 16:55:59 -05:00
Vineeth Voruganti 780bfe1c30
Refactor cache to store data rather than orm objects (#395)
* fix: refactor cache to store data rather than orm objects

* fix: Add Cache Version Keys
2026-02-23 16:03:29 -05:00
Vineeth Voruganti 78df86dc66
fix: Remove noisy sentry error on llm failure and upgrade deps (#396) 2026-02-23 15:53:29 -05:00
ajspig 5bd93a2bef
adding test reasoning levels script (#337)
* chore: 3.0 honcho and 2.0 sdks changelog

fix: use PeerContextResponse in peer.ts

* chore: move docs to /v3/, build SDKs

* chore: code review

* feat: [WIP] migrate away from stainless in typescript sdk

* chore: move api from /v2/ to /v3/

* feat: no-stainless typescript with real tests

* feat: migrate python sdk off of stainless

* feat: clean typescript sdk

* chore: add tests for ts http client

* fix: rewrite entire python sdk in new format, update typescript sdk to use `configuration` not `config` for consistency with API

* fix: clean up SDKs, synchronize

* chore: update sdk examples

* chore: update OpenAPI documentation and SDK examples to reflect changes

* fix: better test

* fix: install deps in test runner, improve robustness of streaming in sdk, coderabbit nits

* fix: standardize around camelCase in TS SDK

* refactor: update configuration handling in SDKs to use typed models for workspace, session, and peer configurations

* docs: clarify queue status usage and remove polling methods from SDKs

add claude skills for migrations

* chore: fix links in docs

* feat: add deriver flush mode to bypass batch token threshold

- Introduced `is_deriver_flush_enabled` function to check if flush mode is active.
- Updated `QueueManager` to conditionally apply batch token thresholds based on flush mode.
- Enhanced `UnifiedTestExecutor` to enable flush mode via Redis.
- Added `flush` parameter to test cases to facilitate testing of flush mode behavior.
- Updated various test cases to utilize the new flush functionality.

* feat: implement schedule_dream functionality in SDKs, use in unified test runner

- Added `schedule_dream` method to both Python and TypeScript SDKs for scheduling dream tasks.
- Updated HTTP routes to include endpoint for scheduling dreams.
- Enhanced test runner to utilize the new `schedule_dream` method for scheduling actions.
- Updated TypeScript client to support the new scheduling functionality with appropriate parameters.

* feat: update single deriver task to support multiple observers

- Changed the `observer` parameter to `observers` as a list in multiple functions across the deriver module.
- Updated the processing logic to handle multiple observers for representation tasks.
- Adjusted related payload and queue management functions to accommodate the new observers structure.
- Modified tests to reflect changes in the representation task handling and ensure proper functionality.

* refactor: update enqueue tests to support deduplication of queue items with multiple observers

- Modified tests in `test_enqueue.py` to reflect changes in the queue item structure, where each message now results in a single queue item containing a list of observers.
- Updated assertions to validate that the `observers` field correctly includes all relevant peers, ensuring proper functionality of the deduplication logic.
- Removed redundant payload matching logic to streamline test cases and improve clarity.

* fix: add backwards compatibility for representation work unit keys and payload observers

* add: results

* add: adoption journey

* feat: update dialectic configuration and introduce cost calculator

- Adjusted LLM and dialectic settings in `.env.template`, `config.toml.example`, and `src/config.py` to reduce maximum tool output characters and session history tokens for cost efficiency.
- Implemented a new `dialectic_cost_calculator.py` script to estimate costs based on reasoning levels and model pricing.
- Enhanced `DialecticAgent` to utilize minimal tools and adjusted output token settings based on reasoning level to optimize performance and reduce costs.

* feat: add reasoning level to chat input in unified test runner

- Enhanced the `UnifiedTestExecutor` to include a `reasoning_level` parameter in the chat method call.
- Updated the `QueryAction` model to support the new `reasoning_level` attribute, allowing for more nuanced chat interactions.

* add: adding script for testing reasoning levels

* fix: moving script

* fix: remove stale doc files deleted in main

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: Code Review Changes

---------

Co-authored-by: Benjamin McCormick <docterformer@protonmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
2026-02-19 15:51:15 -05:00