Skip to main content

Quantization methods catalogue

Hugging Face documents many quantization backends. Use this catalogue as a map — then read the official page for APIs and limitations.

Adapted for this playbook from the 🤗 Transformers documentation by Hugging Face. Official page: https://huggingface.co/docs/transformers/quantization/overview. Images © Hugging Face (documentation-images) unless noted. This is not a substitute for the upstream docs — verify against the current version.

Covers Hugging Face pages: All quantization/* pages listed below

MethodOfficial doc
AQLMaqlm
AutoRoundauto_round
AWQawq
BitNetbitnet
bitsandbytesbitsandbytes
compressed-tensorscompressed_tensors
EETQeetq
FBGEMMfbgemm_fp8
Fine-grained FP8finegrained_fp8
Four Over Sixfouroversix
FP-Quantfp_quant
GGUFgguf
GPTQgptq
HIGGShiggs
HQQhqq
Metalmetal
MXFP4mxfp4
Optimumoptimum
Quantoquanto
Quarkquark
torchaotorchao
SpQRspqr
VPTQvptq
SINQsinq

Contribute improvements via Hugging Face’s quantization contribute guide.

Discussion

Comments​

Share feedback or questions about this page. No account required.

Loading comments…