diff --git a/README.md b/README.md index de70ab1..3113545 100644 --- a/README.md +++ b/README.md @@ -47,6 +47,11 @@ resident run. Opt-in (`stream_layers: true`) and still BETA — [how it works](docs/performance-and-quantization.md#layer-streaming-beta-v0720-nf4-v0722-disk--wider-archs-v0723-preference-losses-v0724) · [all measurements](benchmarks/) · [paper](https://doi.org/10.5281/zenodo.21771064) +
+ 
+ Llama-3.1-8B-Instruct + NF4, LoRA, batch 1, seq 512 on an RTX 3050 Laptop 4 GB — 3.32 GB peak, 119.6 tok/s. Full video (90s)
+