We hold several patents and have published articles, including a CVPR oral presentation on DNNs compression. We've built tools we enjoy using ourselves and hope they can benefit others too!
The framework supports various quantization setups:
- Integer and Float quantization
- Symmetric and Asymmetric quantization
- Dynamic and Static quantization
- Multiple granularity options: per-tensor, per-channel, per-token, etc
- Pre-defined configuration schemas compatible with NVIDIA GPUs that are easy to set up and use