Lossless model compression experiment: GLM-5.2 in 25% less memory(brianbell-x.github.io) |
Lossless model compression experiment: GLM-5.2 in 25% less memory(brianbell-x.github.io) |
Also, are all that many people handling the BF16 weights directly? GLM-5.2's reference deployment is FP8, and many vendors are even serving at NVFP4 which seems to offer negligible degradation over the FP8 reference deployments.