Commit Graph

9 Commits

Author SHA1 Message Date
liyizhouAI 7443082846 feat: v0.3.2 - CustomGraphBuilder replaces Graphiti for step 2 graph build
End of a long debugging rabbit hole. After ~10 attempts patching Graphiti's
compatibility with non-OpenAI LLMs (GLM, Qwen via SiliconFlow), every fix
uncovered another bug. Made the strategic call to abandon Graphiti entirely
for the graph-build path and write a minimal custom builder.

## Why Graphiti had to go

- Qwen 32B 32K context overflow on cumulative episode retrieval (60K-82K tokens)
- SiliconFlow Qwen 72B also 32K (not 128K as docs implied)
- GLM 4 Flash gives 128K context but triggers "20015 parameter invalid" on
  Graphiti's structured output calls (logprobs in reranker, empty input in
  embedder, nested dict in Neo4j writes)
- Each patch target was 1-2 levels deep in Graphiti internals
- 4 separate monkey-patches (EntityNode.save, bulk_utils, driver layer,
  reranker, embedder) still couldn't cover the extract_nodes path

## New architecture

backend/app/services/custom_graph_builder.py (NEW, 348 lines):
- Uses Foresight's own LLMClient (retry + GLM→Qwen fallback + token tracker)
- 1 LLM call per chunk (Graphiti needed 4-5)
- Single prompt extracts entities + relationships as JSON
- Writes directly to Neo4j via cypher MERGE (no Graphiti dependency)
- Schema identical to what Graphiti produced — downstream get_all_nodes /
  get_all_edges / zep_entity_reader all work unchanged
- ThreadPoolExecutor (10 workers) for concurrent chunk extraction
- Sequential Neo4j writes via shared session (avoids name conflicts in dedup)
- DDL wrapped in try/except (tolerates pre-existing Graphiti indexes)
- Entity name dedup via in-memory map → single uuid per canonical name
- Regex-sanitized label / relation type to prevent cypher injection

backend/app/api/graph.py:
- /api/graph/build endpoint now calls CustomGraphBuilder instead of
  builder.add_text_batches() → client.add_episodes_batch() (the old Graphiti path)

backend/app/services/graph_builder.py:
- _build_graph_worker also updated to use CustomGraphBuilder (dead code path
  but kept in sync)

backend/app/models/project.py:
- Default chunk_size 500 → 250 (to keep individual LLM prompts small)

backend/app/services/graphiti_client.py:
- Kept all monkey-patches for backward compat — they now only affect legacy
  Graphiti code paths that CustomGraphBuilder bypasses entirely:
  * _patch_neo4j_driver (AsyncSession.run nested-dict sanitize)
  * _patch_reranker_for_non_openai (GLM logprobs workaround)
  * _patch_embedder_empty_input (SiliconFlow empty-input guard)
  * _patch_entity_node_ops (EntityNode.save sanitize)
  * _patch_add_nodes_and_edges_bulk_tx (bulk-episode sanitize)
- These are kept for safety; Graphiti code path is no longer invoked in the
  production build flow but methods like add_episode still exist on the class

## Performance

- Sequential: ~194 chunks × 2-5s = 10-15 min
- Concurrent (10 workers): ~194 / 10 × 3-5s = **1-2 min** (5-8x speedup)
- Rate limiting handled by LLMClient retry/backoff, not raw thread contention

## Downstream compatibility

Verified:
- get_all_nodes: MATCH (n:Entity) WHERE n.group_id = $gid RETURN n, labels(n)
- get_all_edges: MATCH (a)-[r]->(b) WHERE r.group_id = $gid RETURN r, type(r)
- get_node / get_node_edges: MATCH by uuid

CustomGraphBuilder writes:
- (n:Entity [optional second label]) with uuid, name, summary, group_id, created_at
- [r:RELATION_TYPE] with uuid, name, fact, group_id, created_at

Schema matches exactly.

## Pipeline wiring

Step 1 ontology generation → Step 2 graph build (CustomGraphBuilder) →
Step 3 profile generation (ZepEntityReader queries Neo4j) → Step 4 simulation

No changes needed downstream of Step 2.
2026-04-16 07:54:07 +08:00
liyizhouAI fff7edce2a feat: migrate from Zep Cloud to self-hosted Graphiti + Neo4j
Replace Zep Cloud ($25/mo) with open-source Graphiti + Neo4j:
- New graphiti_client.py: unified wrapper with sync bridge for async Graphiti
- Modified graph_builder.py: use Graphiti add_episode instead of Zep batch API
- Modified zep_entity_reader.py: Neo4j Cypher queries replace Zep pagination
- Modified zep_tools.py: Graphiti search replaces Zep Cloud search
- Modified zep_graph_memory_updater.py: Graphiti add_episode replaces Zep add
- Modified oasis_profile_generator.py: GraphitiClient replaces Zep client
- Updated config.py: NEO4J_URI/USER/PASSWORD replace ZEP_API_KEY
- Updated requirements.txt: graphiti-core + neo4j replace zep-cloud
- Added PRD.md: product requirements document with token analysis

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
2026-04-13 12:07:32 +08:00
liyizhouAI ded715feb2 feat: merge upstream i18n + rebrand MiroFish → Foresight 先见之明
- Merge 41 upstream commits (i18n for 7 languages, security fixes, new features)
- Rebrand all MiroFish references to Foresight/先见之明 across 37 files
- Re-apply dark theme CSS overhaul (pure black/gray, no blue tints)
- Re-apply Teleport-based theme toggle (inline with brand, 20px gap)
- Restore Foresight logo and favicon
- Update GitHub links to liyizhouAI/foresight
- Update locale files, README, package.json

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
2026-04-11 16:40:13 +08:00
ghostubborn 7c07237544 fix(i18n): pass locale to background threads via thread-local storage
Background threads (graph building, simulation prep, report generation,
profile generation) now inherit the requesting user's locale preference.
Previously these fell back to 'zh' because Flask request context was
unavailable in spawned threads.
2026-04-01 16:55:51 +08:00
ghostubborn 9d43b77511 feat(i18n): replace hardcoded Chinese in backend SSE progress messages 2026-04-01 16:32:10 +08:00
666ghj da6548e96f feat(graph): implement pagination for fetching nodes and edges; add utility functions for streamlined data retrieval 2026-02-27 15:53:29 +08:00
666ghj a90b683a44 Enhance graph data retrieval and detail display in Process.vue and graph_builder.py
- Updated the `get_graph_data` method in `graph_builder.py` to include additional attributes such as creation time, validity periods, and episodes for nodes and edges, improving the richness of the graph data.
- Modified the detail panel in `Process.vue` to present new attributes, including properties, episodes, and timestamps, enhancing user interaction and data visibility.
- Improved styling for the detail panel to better organize and present the comprehensive information for selected nodes and edges.
2025-12-10 19:26:30 +08:00
666ghj e98da6b53e Enhance backend startup logging and API endpoint display
- Updated `run.py` to conditionally print startup information only in the reloader process to avoid duplicate logs in debug mode.
- Modified `__init__.py` to log startup and completion messages based on the reloader process condition.
- Added warnings suppression in `graph_builder.py` for Pydantic v2 regarding Field usage.
- Revised `ontology_generator.py` to enforce strict design guidelines for entity types and relationships, ensuring compliance with new requirements.
- Improved logging behavior in `logger.py` to prevent log propagation to the root logger, avoiding duplicate outputs.
2025-11-28 18:59:36 +08:00
666ghj 08f417f3b7 Introduce Project ID for context management, finalizing the stateful API pipeline from file submission to graph construction. 2025-11-28 17:21:08 +08:00