`getChatSuggestions` previously picked the largest installed model by file size, on the assumption that bigger models give better suggestions. This is unsafe: if any installed model exceeds available VRAM (e.g. llama3.1:405b on a 96 GB GPU), Ollama spends minutes trying to load it and the request 500s — making the chat page unusable for anyone who happens to keep a flagship-sized model on disk. Chat suggestions are short prompts that don't benefit from a flagship model anyway. Prefer the user's selected `chat.lastModel` when set, and fall back to the smallest installed model otherwise. `OllamaService.getModels()` already excludes embedders, so the fallback always picks a chat model. |
||
|---|---|---|
| .. | ||
| controllers | ||
| exceptions | ||
| jobs | ||
| middleware | ||
| models | ||
| services | ||
| utils | ||
| validators | ||