The runtime stage copies application code but not pyproject.toml, so
src/_version.py cannot find the file it reads the version from. The
image also installs dependencies with --no-install-project, so there is
no honcho distribution for the importlib.metadata fallback to find.
Both lookups fail, so the service falls back to reporting its version as
"unknown" in the OpenAPI schema and in telemetry events.
Copying the file into the runtime stage restores an accurate version.
The file is under 4 KB, so the image size is unchanged.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The HEALTHCHECK directive probes an HTTP endpoint that only the API
serves. The deriver service reuses this image but is a background queue
worker with no HTTP server — the probe can never succeed, so Docker
permanently marks the deriver container as unhealthy.
Remove the HEALTHCHECK from the shared image. Service-level health
checks belong in each service's own configuration (e.g. Kubernetes
readiness/liveness probes on the API Deployment only).
Closes#521
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: Inconsistencies in Docs, health endpoint, troubleshooting guide
* fix: (docs) maintain consistency on postgres db name
* chore: (docs) update v2 contributing docs with updates db paths
* docs: overhaul self-hosting docs for provider-agnostic setup
- .env.template: lead with provider options (custom, vllm, google,
anthropic, openai, groq) instead of baking in vendor-specific keys.
All provider/model settings commented out so server fails fast until
configured. Separate endpoint config from per-feature provider+model
from tuning knobs.
- docker-compose.yml.example: fix healthcheck -d honcho -> -d postgres
to match POSTGRES_DB=postgres.
- config.toml.example: reorder and document LLM key section with
OpenRouter and vLLM examples.
- self-hosting.mdx: replace multi-vendor key table with provider options
table. Add examples for OpenRouter, vLLM/Ollama, and direct vendor
keys. Remove duplicated key lists from Docker/manual setup sections.
- configuration.mdx: replace scattered provider docs with provider types
table. Fix Docker Compose snippet to match actual compose file. Note
code defaults as fallback, not recommended path.
- troubleshooting.mdx: add alternative provider issues section (custom
provider config, model name format, Docker localhost, structured
output failures).
* docs: add Docker build troubleshooting for permission errors
- Document BuildKit requirement (RUN --mount syntax)
- AppArmor/SELinux blocking Docker builds on Linux
- Volume mount UID mismatch between host and container app user
- Note in self-hosting docs that Docker path builds from source
* docs: reframe self-hosting as contributor/dev path, point to cloud service
* Revert "docs: reframe self-hosting as contributor/dev path, point to cloud service"
This reverts commit 3e766eb1a9.
* docs: add production compose, model guidance, thinking budget docs
- Add docker-compose.prod.yml for VM/server deployment: no source
mounts, restart policies, 127.0.0.1-bound ports, cache enabled
- Add model tier guidance and community quick-start link to self-hosting
- Document THINKING_BUDGET_TOKENS gotcha for non-Anthropic providers
- Add reverse proxy examples (Caddy + nginx) to production section
- Add backup/restore commands to production considerations
* docs: simplify self-hosting to single provider, restructure config guide
Self-hosting page now defaults to one OpenAI-compatible endpoint
with one model for all features. Moved model tiers, alternative
providers, and per-feature tuning into the configuration guide.
Eliminated duplicate config priority sections, dev/prod split,
and redundant TOML examples.
* docs: merge compose files, restore provider/model to feature sections in .env.template
Single docker-compose.yml.example with dev sections commented out.
Moved PROVIDER and MODEL back alongside each feature in .env.template
so settings stay colocated with their module. Updated self-hosting
docs to reference single compose file.
* fix: broken anchor links, redundant migration step, minor inconsistencies
Fix 4 broken internal links (#llm-provider-setup, #llm-api-keys,
#which-api-keys-do-i-need, #alternative-providers) to point to
correct headings. Remove redundant Docker migration step (entrypoint
already runs alembic). Fix cache URL missing ?suppress=true in
reference config. Fix uv install command to use official method.
* docs: env template ready to use, simplify self-hosting flow
.env.template now has provider/model lines uncommented with
placeholder values — user just sets endpoint, key, and model name.
Thinking budgets default to 0 for non-Anthropic providers.
Self-hosting page: removed 30-line env var wall, LLM setup now
points to the template. Merged duplicate verify sections.
Removed api_key from SDK examples (auth off by default).
* docs: reorder next steps, configuration guide first
* fix: default embedding provider to openrouter for single-endpoint setup
Without this, embeddings default to openai which requires a separate
LLM_OPENAI_API_KEY. Setting to openrouter routes embeddings through
the same OpenAI-compatible endpoint as everything else.
* fix: review issues — hermes page, thinking budget, production wording
Hermes integration page: replaced inline Docker/manual setup with
link to self-hosting guide, added elkimek community link. Removed
old env var names (OPENAI_API_KEY without LLM_ prefix).
Troubleshooting: removed "or 1" from thinking budget guidance.
Self-hosting: softened "production-ready" to "production-oriented"
since auth is disabled by default.
* docs: model examples in template, expanded LLM setup, better verify flow
.env.template: added "e.g. google/gemini-2.5-flash" hints next to
model placeholders so users know the expected format.
Self-hosting: expanded LLM Setup to show the 3 things users need to
set (endpoint, key, model name) with find-replace tip. Added build
time note, deriver log check, and real smoke test (create workspace)
to verify section. Health check now notes it doesn't verify DB/LLM.
* fix: smoke test uses v3 API path, not v1
* docs: clarify deriver metrics port vs Prometheus host port
* fix: remove deprecated memoryMode from hermes config example
* docs: update hermes page to match current memory provider config
Updated config to match hermes-agent docs: removed apiKey (not needed
for self-hosted), added hermes memory setup CLI command, added config
fields table (recallMode, writeFrequency, sessionStrategy, etc.).
Better verification tests: store-and-recall across sessions, direct
tool calling test. Links to upstream hermes docs for full field list.
* fix: invalid THINKING_BUDGET_TOKENS=0 and missing docker/ in image
Comment out THINKING_BUDGET_TOKENS=0 in .env.template — deriver,
summary, and dream validators require gt=0. Dialectic levels also
commented out since non-thinking models don't need the override.
Add COPY for docker/ directory in Dockerfile so entrypoint.sh is
available when docker-compose.yml.example references it.
* chore: Additional troubleshooting step
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* fix: add observability to docker compose + get docker compose into a usable state
* fix: init.sql location
* fix(docker): use built venv at runtime
---------
Co-authored-by: adavyas <adavyasharma@gmail.com>
* feat: add CD for honcho images to saas test and prod environments
* fix: use github tag in image label
* feat: split up test and prod deployment flows, push to service after
* fix: action parsing properly hopefully
* fix: proper url, version
* remove excessive fly.toml
* fix: specify prod-image in prod workflow
* Potential fix for code scanning alert no. 12: Workflow does not contain permissions
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Potential fix for code scanning alert no. 11: Workflow does not contain permissions
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Apply suggestions from code review
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
* fix: Change App name in command and make steps sequential
* fix: Address Code Rabbit nitpicks
* Use IMAGE Label environment variable
* fix: collisions between github action groups
* feat: add migrate_db script
* fix: correct image label on prod deploy
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
* chore: Update versioning for release
* fix: remove db creation at start and sync migrations and models
* fix: Checkpoint changing metamessage schema
* chore: linter fixes
* fix: session cloning working
* Hybrid long-term memory (#92)
* Add TOM method switching
* Add system prompt and note on format
* Add persistence tweaks
* Specify format for each section of user representation
* Parse XML tags before saving representation metamessage
* Clean up
* Use Claude 3.5 Haiku and refine prompt
* Simplify message processing
* chore: update token limit on dialectic and model for deriver
* Add embedding-based long-term fact retrieval
* Fix bug preventing new documents from being created
* Use multiple queries + tweak prompt
* Fix collection name bug + add duplicate removal
* First implementation of on-demand user rep generation
* WIP debug on-demand user rep changes
* Fixed representations not being stored & deriver issue
* Some speed improvements
* Play with number of facts / queries
* WIP prompt caching for Claude
* WIP fix anthropic caching
* Anthropic prompt caching working but messages too short
* Use Cerebras for small inferences
* Make dialectic responses 1000 tokens max
* Make user representation generation model a constant
* Use llama 3.1 8b for query generation
* Update env template
* Add crud.get_or_create_protected_collection
* rabbit comments
* Fix linter issues
* Add Cerebras to stream router method
* Better handling of default-empty string args
* Change prints to debug logs
* Add error handling to TOM inference
* Handle missing/empty client in model responses
* Handle no messages case in get_chat_history
* Fix indent
* Add error handling to single_prompt methods
* Fix get_or_create_user_protected_collection
* Simplify openAI-compatible model client instantiation
* Remove health endpoint
* Remove LocalEmbeddingStore
* Change prints to debug logs
* Change sentry track
* Code review changes
* Add README to ToM module
* Switch to Groq
* Fix inconsistent openai compatible provider list in stream()
* Update env template to include Groq variables
* Add model_client tests
* fix: Fix unit tests
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* add scoped API keys (#91)
* add AUTH_JWT_SECRET and ADMIN_KEY, use in security middleware (TODO granular keys)
* WIP: convert all API paths to use scoped keys
* add basic unit tests for API keys, ruff formatting
* MVP of route using JWT for payload
* add get_user_from_token
* add key table to postgres, use it to enable key revocation
* add key revocation pt 2 -- fix order of param checks
* finish convenience routes that assume params from JWT
* add tests for key API
* get_keys
* add secrets utility script, add key rotation, fill out tests
* add tiny cache as PoC
* nits, validations, etc
* only create keys table migration if necessary
* fix keys tests to always use auth
* tiny fix to make custom DATABASE_SCHEMA work
* review: add better docs, fix security issue with cache, clear db on rotation, and more
* remove rotation
* remove key database entirely
* Add `/all` path to get all apps (#94)
* add `/all` path for apps
* assert vector extension installed (need this for groudon)
* review
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* add scoped API keys (#91)
* add AUTH_JWT_SECRET and ADMIN_KEY, use in security middleware (TODO granular keys)
* WIP: convert all API paths to use scoped keys
* add basic unit tests for API keys, ruff formatting
* MVP of route using JWT for payload
* add get_user_from_token
* add key table to postgres, use it to enable key revocation
* add key revocation pt 2 -- fix order of param checks
* finish convenience routes that assume params from JWT
* add tests for key API
* get_keys
* add secrets utility script, add key rotation, fill out tests
* add tiny cache as PoC
* nits, validations, etc
* only create keys table migration if necessary
* fix keys tests to always use auth
* tiny fix to make custom DATABASE_SCHEMA work
* review: add better docs, fix security issue with cache, clear db on rotation, and more
* remove rotation
* remove key database entirely
* Add `/all` path to get all apps (#94)
* add `/all` path for apps
* assert vector extension installed (need this for groudon)
* review
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* chore: README and CHANGELOG updates
* add JWT expiry
* fix: Consolidate get methods with JWT token resolution
* chore: Add Annotation to Path, Query, and Body params
* chore: run ruff formatter
* chore: nits & add one exhaustive test of a query route
* fix: undo change to fly.toml
* fix: Langfuse tracing
* Consolidate Get Methods (#96)
* fix: Consolidate get methods with JWT token resolution
* chore: Add Annotation to Path, Query, and Body params
* chore: run ruff formatter
* chore: nits & add one exhaustive test of a query route
* fix: undo change to fly.toml
---------
Co-authored-by: dr-frmr <docterformer@protonmail.com>
* fix: dev-667 fix streaming endpoint
* fix: Anthropic Langfuse Tracing
* fix: add scripts folder to dockerfile
* fix: Remove redundant fields from pydantic schemas
* fix: Add deeper protection on reserved collection
* fix: Consolidate chat and stream methods
* docs: Update Mintlify API Reference and Changelog
* remove langchain guide, update architecture diagram
* honcho mcp server
* chore: Update .env template
* update discord, temporarily remove other guides
* Limit dialectic & deriver context usage with two-scale progressive summarization (#97)
* WIP two tiered summaries
* Move to process_item
* Save user rep metamessage even if no message_id
* Change number of messages per short summary
* Fix broken mock
* Remove prints
* chore: fix test
---------
Co-authored-by: Vineeth Voruganti <13438633+VVoruganti@users.noreply.github.com>
* feat: Add Gemini Support, link facts to message, use 8b for dialectic fact queries
* chore: Styling
* chore: coderabbit nitpicks
* keep dialectic guide
* Add streaming guide
* Remove TODO from dialectic guide
* Fix JS snippets that referred to honcho singleton as client
* Add App explanation to architecture page
---------
Co-authored-by: Dani Balcells <18307962+danibalcells@users.noreply.github.com>
Co-authored-by: doria <93405247+dr-frmr@users.noreply.github.com>
Co-authored-by: dr-frmr <docterformer@protonmail.com>
Co-authored-by: vintro <vince@plasticlabs.ai>
Co-authored-by: Daniel Balcells <dbalcells@gmail.com>
* feat(dialectic) Allow for batch questions and load session history
* feat(dialectic) parallelize facts and history queries
* feat(agent) Addresses dev-258 allow specifying additional collections
* feat(uv) switched from poetry to uv
* feat(deriver) Turn off derivations by editing session medatadata with a deriver_disabled flag
* fix(tests) Clean up test logic
---------
Co-authored-by: Vineeth Voruganti <vineeth@macbook-pro.mynetworksettings.com>