[演講訊息]12/18(三)Inference Accelerators: Use Term Ranking Rather Than Quantization 主講人:Professor H.T. Kung (孔祥重教授)
國立清華大學資訊工程學系
Department of Computer Science
National Tsing Hua University
專題演講
SEMINAR
主講人:Professor H.T. Kung (孔祥重教授),
SPEAKER (Harvard University)
題 目:Inference Accelerators: Use Term Ranking Rather
TOPIC Than Quantization
時 間:108年12月18日(三)下午1點30分至3點
DATE
地 點:資電館地演廳
PLACE
Abstract:
In designing accelerators for convolutional neural network (CNN) inference, we can achieve further compression (e.g., 2~3x) beyond conventional quantization and pruning. We introduce “term revealing,” a procedure that selects top power-of-two terms to use, based on term ranking, for groups of dot-product computations. Due to term revealing’s use of nonlinear term ranking, truncation has significantly less impact on both quantization error and classification accuracy compared to quantization. We have implemented term revealing with FPGA and ASIC designs for 8-bit CNN models pre-trained on ImageNet, such as ResNet-152 (no model retraining is required). These designs show significant gains in latency and energy efficiency while allowing less than 0.5% decrease in top-1 accuracy in ImageNet. (Joint work with Brad McDanel and Sai Zhang at Harvard.)
Bio:
H. T. Kung is William H. Gates Professor of Computer Science and Electrical Engineering at Harvard University. He has pursued a variety of research interests in his career, ranging from computer science theory, parallel computing, VLSI design, database algorithms, computer systems, wireless communications, and networking, to machine learning. His academic honors include Member of Academia Sinica (Taiwan), Member of National Academy of Engineering (USA), Guggenheim Fellowship, and the ACM SIGOPS 2015 Hall of Fame Award (with John Robinson). He currently serves as volunteer President of a non-profit organization, Taiwan AI Academy, with campuses in four cities nurturing thousands of AI talents for the industry.
--
