Dense source content produces chunks that exceed the embedding model's
context window (nomic-embed-text:v1.5 defaults to 2048 tokens). Two paths
hit this even after the prior pre-cap:
- Older Ollama (e.g. 0.18.1, #944) ignores the num_ctx=8192 we send on
/api/embed, so it stays at the model's 2048 default.
- The OpenAI-compat /v1/embeddings fallback didn't pass num_ctx/truncate
at all, so any Ollama drops to 2048 whenever it lands on the fallback.
When a chunk overflowed, the 400 was swallowed and the chunk was silently
dropped from Qdrant. Worse, the failure propagated to EmbedFileJob, which
re-embeds the entire file on each of its 30 BullMQ attempts — the "endless
queue loop" / "api/embed for weeks" / pegged GPU reported in #944/#959.
Fix:
- OllamaService.embed(): on a context-length error, retry once with an
aggressive 2048-safe cap (EMBED_CONTEXT_SAFE_CHARS = 2000) so the chunk
is embedded (start-of-chunk) instead of dropped. Native-path context
errors now bubble to this retry instead of falling through to the
smaller-context fallback. Split the native+fallback attempt into
_embedWithFallback().
- Pass truncate/num_ctx on the /v1/embeddings fallback too (Ollama's
OpenAI-compat shim forwards them).
- EmbedFileJob: classify "input length exceeds context length" as an
UnrecoverableError so one permanently-oversized chunk can't trigger 30
full-file re-embeds.
- Add OllamaService.isContextLengthError() shared by both.
Graceful degradation: a truncated chunk loses its tail but is kept in the
index, which is strictly better than today's silent drop + retry storm.
Refs #881. Supersedes the #369/#670 symptom closures that never fixed the
fallback path.