Why FastForward#
-
It is still PyTorch
A
Modulegoes in, aModulecomes out. FastForward extends the PyTorch dispatcher, so your training loop, your checkpoints and your optimizer do not change. -
Made for fast iterations
find_quantizers(…)reaches any layer with a selector,.initialize(…)sets bitwidth and granularity from the outside. Change the numbers, run again, you never edit the model. -
Extensible
Quantizers, range estimators, and operators are all plug-in points. Ideal for research on new quantization methods.
Features#
-
Quantized Tensor
A versatile container for quantized data that supports multiple quantization formats while retaining metadata.
-
Range Estimation
General methods for range estimation, easy to extend to new quantization schemes.
-
Operator Dispatch
A dispatcher built on top of PyTorch's, specialized for different quantization schemes and methods.
-
Safe Quantization
Ensures all the executed operations are quantized, with explicit exceptions if needed. This helps catch common quantization mistakes early.
-
MPath
Describe, access, and update layers deep in a module hierarchy at a higher level of abstraction. Tutorial.
-
Autoquant (experimental)
Automatic conversion of any PyTorch model into an eager-mode quantized model. Tutorial.
Quick look#
import fastforward as ff
import torch
# 1. Convert any PyTorch model to a quantization-ready one.
model = MyModel()
ff.quantize_model(model)
# 2. Attach 8-bit per-channel weight quantizers to every linear layer.
weight_quantizers = ff.find_quantizers(model, "**/[quantizer:parameter/weight]")
weight_quantizers.initialize(ff.nn.LinearQuantizer, num_bits=8, granularity=ff.PerChannel())
# 3. Calibrate on real data.
with ff.estimate_ranges(model, ff.range_setting.RunningMinMaxRangeEstimator):
for batch in calibration_loader:
model(**batch)
# 4. Run the quantized model like any PyTorch model — pdb and print still work.
output = model(**input_batch)
Full walkthrough on Llama-v3 All tutorials
Install#
Requires a working PyTorch install (≥ 2.4).
pip install git+https://github.com/Qualcomm-AI-research/fastforward@main
Citation#
If you use FastForward in your research, please cite it as:
@software{fastforward,
title = {FastForward: A PyTorch-based Library for Neural Network Quantization},
author = {Peters, Jorn and Behrends, S{\"o}nke and Del Chiaro, Riccardo and
Mironov, Evgeny and van Rozendaal, Ties and Stasis, Spyridon and
Weitkamp, Laurens and Nagel, Markus},
year = {2024},
url = {https://github.com/Qualcomm-AI-research/fastforward}
}
Project status
FastForward is under active development. It is already used in research and production projects at Qualcomm AI Research, but core APIs may still evolve. Roadmap items include additional post-training methods and richer export targets.
