Do LLM quants impact as much as we think?(rakuensoftware.com) |
Do LLM quants impact as much as we think?(rakuensoftware.com) |
I've had my 5080 running 24/7 over the past 2 weeks to try to understand the impact of quants on models. Came to some strange conclusions which were interesting to me.
I didn't really see a decline until going under Q4 for the most part, and MoEs were impacted far less then dense models.
Kind of cool results, and unexpected. Although this is just an initial set of tests to validate if the quant itself damaged the model, I've got a longer article coming in the future that's going to test these models in devops, coding, and long sessions. Also, I'm aware it's only a 16GiB card! I've got another article series coming out in the near future for much larger VRAM!
The benchmark code, dataset, and results are open sourced. Take a look, tell me I'm wrong (Wouldn't be the first time!), or run the benchmarks yourself!
There’s a parallel in quantitative trading. It’s very easy to build a hierarchy from theory — more precision, more parameters, a better Sharpe, a more sophisticated model — and then unconsciously treat that hierarchy as an empirical fact.
The interesting question is usually: under the conditions we actually care about, can we reliably distinguish A from B?
I’d be interested to see whether the separation between the quants becomes clearer as session length and task complexity increase. That feels like the equivalent of moving a trading model from a clean backtest into progressively less forgiving production conditions (which is kinda what I'm obsessed with).