* fix(rag): add proper .docx text extraction via mammoth .docx files were classified as plain text and routed through raw-text extraction (extractTXTText). Since .docx is a ZIP archive containing XML, this produced garbage content in the Knowledge Base — XML tags and binary noise instead of the actual document text. Adds a dedicated 'docx' file type (split out of the generic 'text' bucket in determineFileType) and a mammoth-based extractor that parses the document XML properly. * fix: clean up package-lock.json diff * chore(deps): pin mammoth version --------- Co-authored-by: John Cortright <jcortright@zscaler.com> Co-authored-by: jakeaturner <jturner@cosmistack.com> |
||
|---|---|---|
| .. | ||
| app_auto_update_service.ts | ||
| auto_update_service.ts | ||
| benchmark_service.ts | ||
| benchmark_telemetry.ts | ||
| chat_service.ts | ||
| collection_manifest_service.ts | ||
| collection_update_service.ts | ||
| container_registry_service.ts | ||
| content_auto_update_service.ts | ||
| countries_service.ts | ||
| custom_app_guard.ts | ||
| docker_service.ts | ||
| docs_service.ts | ||
| download_service.ts | ||
| kiwix_catalog_service.ts | ||
| kiwix_library_service.ts | ||
| map_service.ts | ||
| ollama_service.ts | ||
| queue_service.ts | ||
| rag_service.ts | ||
| system_service.ts | ||
| system_update_service.ts | ||
| zim_extraction_service.ts | ||
| zim_service.ts | ||