Deep learning has been successfully applied to a variety of tasks. On real-time scenarios such as inference on autonomous vehicles, the inference speed of the model is critical. Network quantization is an effective approach to accelerating deep learning models.
Automating Optimization of Quantized Deep Learning Models on CUDA
calendar_today
April 29, 2019
domain
apache-tvm