Collect the state-dict pieces emitted by each QuantizedTensor weight instead of filtering globally by suffix. This preserves real module parameters such as input_scale, keeps shared module aliases working, and leaves the real weight dequantization path intact. Signed-off-by: liminfei-amd <91481003+liminfei-amd@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| test_mixed_precision.py | ||