Projects by Tag: Quantization — Model Compression, INT8 & Quantization-Aware Training for Efficient Edge Inference
Explore projects tagged with quantization to discover neural network model compression implementations—covering post-training quantization (PTQ), quantization-aware training (QAT), INT8/INT4 workflows, and mixed-precision strategies for edge and mobile deployment. This curated list of projects highlights real-world benchmarks, accuracy-preserving optimization tactics, compatible toolchains (ONNX, TensorRT, TFLite), and deployment guides to reduce model size, lower latency, and improve power efficiency. Use the filtering UI above to narrow results by framework, precision, dataset, or deployment target, then click through to repos and step-by-step guides for reproducible quantization pipelines; start comparing trade-offs, integrating quantization into CI/CD, and deploying optimized models to production.