Keep the transformer and Qwen text encoder off CUDA during initial load/quantization in low-VRAM mode so model startup avoids full-model OOM before offloading and quantization can take effect. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Jaret Burkett <jaretburkett@gmail.com> |
||
|---|---|---|
| .. | ||
| src | ||
| __init__.py | ||
| flux2_klein_model.py | ||
| flux2_model.py | ||